Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks.
Charlotte Siska, Katerina Marazopoulou, Melissa Ailem, James Bono
Browse the full ACL paper archive.
Charlotte Siska, Katerina Marazopoulou, Melissa Ailem, James Bono
Browse the full ACL paper archive.