Skip to content

LAD: Efficient Accelerator for Generative Inference of LLM with Locality Aware Decoding.

Haoran Wang, Yuming Li, Haobo Xu, Ying Wang, Liqi Liu, Jun Yang, Yinhe Han

VenueA*HPCA
Year2025
ProceedingsHPCA

Browse the full HPCA paper archive.