Optimality and Approximation with Policy Gradient Methods in Markov Decision Processes.
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, Gaurav Mahajan
Browse the full COLT paper archive.
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, Gaurav Mahajan
Browse the full COLT paper archive.