Skip to content

SplitQuant: Resource-Efficient LLM Offline Serving on Heterogeneous GPUs via Phase-Aware Model Partition and Adaptive Quantization.

Juntao Zhao, Borui Wan, Yanghua Peng, Haibin Lin, Chuan Wu

Year2025
ProceedingsCLUSTER

Browse the full CLUSTER paper archive.