Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference.
Hao Mark Chen, Wayne Luk, Yiu Ka Fai Cedric, Rui Li, Konstantin Mishchenko, Stylianos I. Venieris, Hongxiang Fan
Browse the full EMNLP paper archive.
Hao Mark Chen, Wayne Luk, Yiu Ka Fai Cedric, Rui Li, Konstantin Mishchenko, Stylianos I. Venieris, Hongxiang Fan
Browse the full EMNLP paper archive.