Skip to content

SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration.

Heming Xia, Yongqi Li, Jun Zhang, Cunxiao Du, Wenjie Li

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.