GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference.
Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, Andreas Moshovos
Browse the full MICRO paper archive.
Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, Andreas Moshovos
Browse the full MICRO paper archive.