Skip to content

Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP.

Yuanhao Wang, Kefan Dong, Xiaoyu Chen, Liwei Wang

VenueA*ICLR
Year2020
ProceedingsICLR

Browse the full ICLR paper archive.