Variance-Dependent Regret Bounds for Linear Bandits and Reinforcement Learning: Adaptivity and Computational Efficiency.
Heyang Zhao, Jiafan He, Dongruo Zhou, Tong Zhang, Quanquan Gu
Browse the full COLT paper archive.
Heyang Zhao, Jiafan He, Dongruo Zhou, Tong Zhang, Quanquan Gu
Browse the full COLT paper archive.