Minimax Optimal Regret Bound for Reinforcement Learning with Trajectory Feedback.
Zihan Zhang, Yuxin Chen, Jason D. Lee, Simon Shaolei Du, Ruosong Wang
Browse the full ICML paper archive.
Zihan Zhang, Yuxin Chen, Jason D. Lee, Simon Shaolei Du, Ruosong Wang
Browse the full ICML paper archive.