Skip to content

Nearly Minimax Optimal Regret for Learning Infinite-horizon Average-reward MDPs with Linear Function Approximation.

Yue Wu, Dongruo Zhou, Quanquan Gu

Year2022
ProceedingsAISTATS

Browse the full AISTATS paper archive.