Learning Infinite-horizon Average-reward MDPs with Linear Function Approximation.
Chen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, Rahul Jain
Browse the full AISTATS paper archive.
Chen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, Rahul Jain
Browse the full AISTATS paper archive.