Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective.
Dylan J. Foster, Alexander Rakhlin, David Simchi-Levi, Yunzong Xu
Browse the full COLT paper archive.
Dylan J. Foster, Alexander Rakhlin, David Simchi-Levi, Yunzong Xu
Browse the full COLT paper archive.