Horizon-Free and Instance-Dependent Regret Bounds for Reinforcement Learning with General Function Approximation.
Jiayi Huang, Han Zhong, Liwei Wang, Lin Yang
Browse the full AISTATS paper archive.
Jiayi Huang, Han Zhong, Liwei Wang, Lin Yang
Browse the full AISTATS paper archive.