Skip to content

Federated Q-Learning with Reference-Advantage Decomposition: Almost Optimal Regret and Logarithmic Communication Cost.

Zhong Zheng, Haochen Zhang, Lingzhou Xue

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.