Federated Q-Learning with Reference-Advantage Decomposition: Almost Optimal Regret and Logarithmic Communication Cost.
Zhong Zheng, Haochen Zhang, Lingzhou Xue
Browse the full ICLR paper archive.
Zhong Zheng, Haochen Zhang, Lingzhou Xue
Browse the full ICLR paper archive.