Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning.
Qiwei Di, Heyang Zhao, Jiafan He, Quanquan Gu
Browse the full ICLR paper archive.
Qiwei Di, Heyang Zhao, Jiafan He, Quanquan Gu
Browse the full ICLR paper archive.