Throughput-Oriented LLM Inference via KV-Activation Hybrid Caching with A Single GPU.
Sanghyeon Lee, Hongbeen Kim, Soojin Hwang, Guseul Heo, Minwoo Noh, Jaehyuk Huh
Browse the full ICCD paper archive.
Sanghyeon Lee, Hongbeen Kim, Soojin Hwang, Guseul Heo, Minwoo Noh, Jaehyuk Huh
Browse the full ICCD paper archive.