Skip to content

It Takes Two: Embracing Sparsity and Speculative Decoding for Efficient LLM Inference.

Haolin Chu, Changyu Chen, Wei Liu, Jian Luan, Jiabin Deng, Liang Liu, Huadong Ma, Xiaolong Zheng

VenueBIWQoS
Year2026
ProceedingsIWQoS

Browse the full IWQoS paper archive.