Skip to content

WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

Bill Yuchen Lin, Yuntian Deng, Khyathi Raghavi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, Yejin Choi

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.