On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, Ping Tak Peter Tang
Browse the full ICLR paper archive.
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, Ping Tak Peter Tang
Browse the full ICLR paper archive.