Nearly Minimax Optimal Regret for Learning Infinite-horizon Average-reward MDPs with Linear Function Approximation.
Yue Wu, Dongruo Zhou, Quanquan Gu
Browse the full AISTATS paper archive.
Yue Wu, Dongruo Zhou, Quanquan Gu
Browse the full AISTATS paper archive.