Reward Estimation for Variance Reduction in Deep Reinforcement Learning.
Joshua Romoff, Alexandre Pich, Peter Henderson, Vincent Franois-Lavet, Joelle Pineau
Browse the full ICLR paper archive.
Joshua Romoff, Alexandre Pich, Peter Henderson, Vincent Franois-Lavet, Joelle Pineau
Browse the full ICLR paper archive.