A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima.
Zeke Xie, Issei Sato, Masashi Sugiyama
Browse the full ICLR paper archive.
Zeke Xie, Issei Sato, Masashi Sugiyama
Browse the full ICLR paper archive.