Skip to content

RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression.

Payman Behnam, Yaosheng Fu, Ritchie Zhao, Po-An Tsai, Zhiding Yu, Alexey Tumanov

VenueA*ICML
Year2025
ProceedingsICML

Browse the full ICML paper archive.