WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.
Bill Yuchen Lin, Yuntian Deng, Khyathi Raghavi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, Yejin Choi
Browse the full ICLR paper archive.