Skip to content

SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget.

Zihao Wang, Bin Cui, Shaoduo Gan

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.