Skip to content

Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference.

Harry Dong, Xinyu Yang, Zhenyu Zhang, Zhangyang Wang, Yuejie Chi, Beidi Chen

VenueA*ICML
Year2024
ProceedingsICML

Browse the full ICML paper archive.