Optimal Regret Bounds for Selecting the State Representation in Reinforcement Learning.
Odalric-Ambrym Maillard, Phuong Nguyen, Ronald Ortner, Daniil Ryabko
Browse the full ICML paper archive.
Odalric-Ambrym Maillard, Phuong Nguyen, Ronald Ortner, Daniil Ryabko
Browse the full ICML paper archive.