XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference.
Joo Monteiro, tienne Marcotte, Pierre-Andr Nol, Valentina Zantedeschi, David Vzquez, Nicolas Chapados, Christopher Pal, Perouz Taslakian
Browse the full EMNLP paper archive.
Joo Monteiro, tienne Marcotte, Pierre-Andr Nol, Valentina Zantedeschi, David Vzquez, Nicolas Chapados, Christopher Pal, Perouz Taslakian
Browse the full EMNLP paper archive.