Skip to content

Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers.

Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, Joey Gonzalez

VenueA*ICML
Year2020
ProceedingsICML

Browse the full ICML paper archive.