Skip to content

InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference.

Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, Jie Zhang

VenueA*HPCA
Year2025
ProceedingsHPCA

Browse the full HPCA paper archive.