Skip to content

Variance-Dependent Regret Bounds for Linear Bandits and Reinforcement Learning: Adaptivity and Computational Efficiency.

Heyang Zhao, Jiafan He, Dongruo Zhou, Tong Zhang, Quanquan Gu

VenueA*COLT
Year2023
ProceedingsCOLT

Browse the full COLT paper archive.