Parallel Restarted SGD with Faster Convergence and Less Communication: Demystifying Why Model Averaging Works for Deep Learning.
Hao Yu, Sen Yang, Shenghuo Zhu
Browse the full AAAI paper archive.
Hao Yu, Sen Yang, Shenghuo Zhu
Browse the full AAAI paper archive.