Skip to content

Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?

Cheng Zhang, Jianyi Cheng, Ilia Shumailov, George A. Constantinides, Yiren Zhao

VenueA*EMNLP
Year2023
ProceedingsEMNLP

Browse the full EMNLP paper archive.