Logarithmic Regret for Online KL-Regularized Reinforcement Learning.
Heyang Zhao, Chenlu Ye, Wei Xiong, Quanquan Gu, Tong Zhang
Browse the full ICML paper archive.
Heyang Zhao, Chenlu Ye, Wei Xiong, Quanquan Gu, Tong Zhang
Browse the full ICML paper archive.