Cache-Efficient Posterior Sampling for Reinforcement Learning with LLM-Derived Priors Across Discrete and Continuous Domains.
Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
Browse the full EMNLP paper archive.
Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
Browse the full EMNLP paper archive.