Is Reinforcement Learning More Difficult Than Bandits? A Near-optimal Algorithm Escaping the Curse of Horizon.
Zihan Zhang, Xiangyang Ji, Simon S. Du
Browse the full COLT paper archive.
Zihan Zhang, Xiangyang Ji, Simon S. Du
Browse the full COLT paper archive.