Skip to content

Layer-Condensed KV Cache for Efficient Inference of Large Language Models.

Haoyi Wu, Kewei Tu

VenueA*ACL
Year2024
ProceedingsACL (1)

Browse the full ACL paper archive.