Skip to content

Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the data.

Florian E. Dorner, Vivian Yvonne Nastl, Moritz Hardt

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.