SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration.
Heming Xia, Yongqi Li, Jun Zhang, Cunxiao Du, Wenjie Li
Browse the full ICLR paper archive.
Heming Xia, Yongqi Li, Jun Zhang, Cunxiao Du, Wenjie Li
Browse the full ICLR paper archive.