Skip to content

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators.

Yilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong, Shafiq Joty

VenueA*ICML
Year2025
ProceedingsICML

Browse the full ICML paper archive.