Train Flat, Then Compress: Sharpness-Aware Minimization Learns More Compressible Models.
Clara Na, Sanket Vaibhav Mehta, Emma Strubell
Browse the full EMNLP paper archive.
Clara Na, Sanket Vaibhav Mehta, Emma Strubell
Browse the full EMNLP paper archive.