Skip to content

Online Target Q-learning with Reverse Experience Replay: Efficiently finding the Optimal Policy for Linear MDPs.

Naman Agarwal, Syomantak Chaudhuri, Prateek Jain, Dheeraj Mysore Nagaraj, Praneeth Netrapalli

VenueA*ICLR
Year2022
ProceedingsICLR

Browse the full ICLR paper archive.