Skip to content

Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Yang You, Jing Li, Sashank J. Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, Cho-Jui Hsieh

VenueA*ICLR
Year2020
ProceedingsICLR

Browse the full ICLR paper archive.