Data-dependent Bounds with T-Optimal Best-of-Both-Worlds Guarantees in Multi-Armed Bandits using Stability-Penalty Matching.
Quan M. Nguyen, Shinji Ito, Junpei Komiyama, Nishant A. Mehta
Browse the full COLT paper archive.
Quan M. Nguyen, Shinji Ito, Junpei Komiyama, Nishant A. Mehta
Browse the full COLT paper archive.