Skip to content

Scaling LLM Inference Efficiently with Optimized Sample Compute Allocation.

Kexun Zhang, Shang Zhou, Danqing Wang, William Yang Wang, Lei Li

VenueANAACL
Year2025
ProceedingsNAACL (Long Papers)

Browse the full NAACL paper archive.