Skip to content

A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima.

Zeke Xie, Issei Sato, Masashi Sugiyama

VenueA*ICLR
Year2021
ProceedingsICLR

Browse the full ICLR paper archive.