Skip to content

GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference.

Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, Andreas Moshovos

VenueA*MICRO
Year2020
ProceedingsMICRO

Browse the full MICRO paper archive.