Skip to content

Throughput-Oriented LLM Inference via KV-Activation Hybrid Caching with A Single GPU.

Sanghyeon Lee, Hongbeen Kim, Soojin Hwang, Guseul Heo, Minwoo Noh, Jaehyuk Huh

VenueCICCD
Year2025
ProceedingsICCD

Browse the full ICCD paper archive.