Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores.
Shaobo Ma, Chao Fang, Haikuo Shao, Zhongfeng Wang
Browse the full ASPDAC paper archive.
Shaobo Ma, Chao Fang, Haikuo Shao, Zhongfeng Wang
Browse the full ASPDAC paper archive.