Skip to content

Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference.

Hao Mark Chen, Wayne Luk, Yiu Ka Fai Cedric, Rui Li, Konstantin Mishchenko, Stylianos I. Venieris, Hongxiang Fan

VenueA*EMNLP
Year2025
ProceedingsEMNLP (Findings)

Browse the full EMNLP paper archive.