Pushing Bit-Width Limits in LLM Quantization with Saliency-Guided Mix-Precision Allocation and Learnable Affine Transformation.
Shuoyu Ma, Wenrui Dai, Maida Cao, Shaohui Li, Ziyang Zheng, Chenglin Li, Junni Zou, Hongkai Xiong
Browse the full DCC paper archive.