Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span.
Woojin Chae, Kihyuk Hong, Yufan Zhang, Ambuj Tewari, Dabeen Lee
Browse the full AISTATS paper archive.
Woojin Chae, Kihyuk Hong, Yufan Zhang, Ambuj Tewari, Dabeen Lee
Browse the full AISTATS paper archive.