Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies.
Nadav Timor, Jonathan Mamou, Daniel Korat, Moshe Berchansky, Gaurav Jain, Oren Pereg, Moshe Wasserblat, David Harel
Browse the full ICML paper archive.