Skip to content

SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading.

Tu Anh Dinh, Carlos Mullov, Leonard Brmann, Zhaolin Li, Danni Liu, Simon Rei, Jueun Lee, Nathan Lerzer, Jianfeng Gao, Fabian Peller-Konrad, Tobias Rddiger, Alexander Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Bhm, Jan Niehues

VenueA*EMNLP
Year2024
ProceedingsEMNLP

Browse the full EMNLP paper archive.