Skip to content

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency.

Ruixiao Li, Fahao Chen, Peng Li

VenueA*IJCAI
Year2025
ProceedingsIJCAI

Browse the full IJCAI paper archive.