Skip to content

When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards.

Norah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yousef Almushayqih, Faisal Mirza, Nouf Alotaibi, Nora Al-Twairesh, Areeb Alowisheq, M. Saiful Bari, Haidar Khan

VenueA*ACL
Year2024
ProceedingsACL (1)

Browse the full ACL paper archive.