Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation.
Jianliang He, Han Zhong, Zhuoran Yang
Browse the full ICLR paper archive.
Jianliang He, Han Zhong, Zhuoran Yang
Browse the full ICLR paper archive.