The Mirage of Action-Dependent Baselines in Reinforcement Learning.
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard E. Turner, Zoubin Ghahramani, Sergey Levine
Browse the full ICLR paper archive.
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard E. Turner, Zoubin Ghahramani, Sergey Levine
Browse the full ICLR paper archive.