Skip to content

Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, Zhifang Sui

VenueA*ACL
Year2024
ProceedingsACL (Findings)

Browse the full ACL paper archive.