Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization.
Seungwoo Son, Wonpyo Park, Woohyun Han, Kyuyeun Kim, Jaeho Lee
Browse the full EMNLP paper archive.
Seungwoo Son, Wonpyo Park, Woohyun Han, Kyuyeun Kim, Jaeho Lee
Browse the full EMNLP paper archive.