Cache Saver: A Modular Framework for Efficient, Affordable, and Reproducible LLM Inference.
Nearchos Potamitis, Lars Henning Klein, Bardia Mohammadi, Chongyang Xu, Attreyee Mukherjee, Niket Tandon, Laurent Bindschaedler, Akhil Arora
Browse the full EMNLP paper archive.