Stop Regressing: Training Value Functions via Classification for Scalable Deep RL.
Jesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taga, Yevgen Chebotar, Ted Xiao, Alex Irpan, Sergey Levine, Pablo Samuel Castro, Aleksandra Faust, Aviral Kumar, Rishabh Agarwal
Browse the full ICML paper archive.