ByteScale: Communication-Efficient Scaling of LLM Training with a 2048K Context Length on 16384 GPUs.
Hao Ge, Junda Feng, Qi Huang, Fangcheng Fu, Xiaonan Nie, Lei Zuo, Haibin Lin, Bin Cui, Xin Liu
Browse the full SIGCOMM paper archive.
Hao Ge, Junda Feng, Qi Huang, Fangcheng Fu, Xiaonan Nie, Lei Zuo, Haibin Lin, Bin Cui, Xin Liu
Browse the full SIGCOMM paper archive.