Skip to content

Serving Long-Context LLMs at the Mobile Edge: Test-Time Reinforcement Learning-based Model Caching and Inference Offloading.

Minrui Xu, Dusit Niyato, Christopher G. Brinton

Year2025
ProceedingsGLOBECOM

Browse the full GLOBECOM paper archive.