Skip to content

Optimizing GPU Multiplexing for Efficient and Cost-Effective Access to Diverse Large Language Models in GPU Clusters.

Yue Zhu, Chen Wang, Max Calman, Rina Nakazawa, Eun Kyung Lee

Year2024
ProceedingsMASCOTS

Browse the full MASCOTS paper archive.