| 2026 | AIED | How Linguistic Diversity Impacts Multilingual Automated Scoring in Large-Scale Assessments. | Hyo Jeong Shin, Nico Andersen, Euigyum Kim, Andrea Horbach, Fabian Zehner |
| 2026 | LREC | GerVLPro: A CEFR-Graded Vocabulary List of L2 Learners' Productive Vocabulary in German. | Noah-Manuel Michael, Anna Hlsing, Andrea Horbach |
| 2026 | LREC | GENIUS Keylog Corpus - a German High School Student Corpus with Keystroke Logging Data. | Nils-Jonathan Schaller, Thorben Jansen, Lars Hft, Hannah Pnjer, Andrea Horbach |
| 2025 | COLING | FEAT-writing: An Interactive Training System for Argumentative Writing. | Yuning Ding, Franziska Wehrhahn, Andrea Horbach |
| 2025 | LAK | Feedback from Generative AI: Correlates of Student Engagement in Text Revision from 655 Classes from Primary and Secondary School. | Thorben Jansen, Andrea Horbach, Jennifer Meyer |
| 2024 | COLING | EVil-Probe - a Composite Benchmark for Extensive Visio-Linguistic Probing. | Marie Bexte, Andrea Horbach, Torsten Zesch |
| 2024 | COLING | When Argumentation Meets Cohesion: Enhancing Automatic Feedback in Student Writing. | Yuning Ding, Omid Kashefi, Swapna Somasundaran, Andrea Horbach |
| 2024 | COLING | DARIUS: A Comprehensive Learner Corpus for Argument Mining in German-Language Essays. | Nils-Jonathan Schaller, Andrea Horbach, Lars Ingver Hft, Yuning Ding, Jan Luca Bahr, Jennifer Meyer, Thorben Jansen |
| 2024 | EACL | Rainbow - A Benchmark for Systematic Testing of How Sensitive Visio-Linguistic Models are to Color Naming. | Marie Bexte, Andrea Horbach, Torsten Zesch |
| 2023 | ACL | Similarity-Based Content Scoring - A more Classroom-Suitable Alternative to Instance-Based Scoring? | Marie Bexte, Andrea Horbach, Torsten Zesch |
| 2023 | ACL | Score It All Together: A Multi-Task Learning Study on Automatic Scoring of Argumentative Essays. | Yuning Ding, Marie Bexte, Andrea Horbach |
| 2022 | LREC | LeSpell - A Multi-Lingual Benchmark Corpus of Spelling Errors to Develop Spellchecking Methods for Learner Language. | Marie Bexte, Ronja Laarmann-Quante, Andrea Horbach, Torsten Zesch |
| 2020 | COLING | Don't take "nswvtnvakgxpm" for an answer -The surprising vulnerability of automatic content scoring systems to adversarial input. | Yuning Ding, Brian Riordan, Andrea Horbach, Aoife Cahill, Torsten Zesch |
| 2020 | IJCNLP | Chinese Content Scoring: Open-Access Datasets and Features on Different Segmentation Levels. | Yuning Ding, Andrea Horbach, Torsten Zesch |
| 2020 | LREC | Linguistic Appropriateness and Pedagogic Usefulness of Reading Comprehension Questions. | Andrea Horbach, Itziar Aldabe, Marie Bexte, Oier Lopez de Lacalle, Montse Maritxalar |
| 2018 | LREC | Semi-Supervised Clustering for Short Answer Scoring. | Andrea Horbach, Manfred Pinkal |
| 2018 | LREC | ESCRITO - An NLP-Enhanced Educational Scoring Toolkit. | Torsten Zesch, Andrea Horbach |
| 2016 | LREC | Unsupervised Ranked Cross-Lingual Lexical Substitution for Low-Resource Languages. | Stefan Ecker, Andrea Horbach, Stefan Thater |
| 2016 | LREC | A Corpus of Literal and Idiomatic Uses of German Infinitive-Verb Compounds. | Andrea Horbach, Andrea Hensler, Sabine Krome, Jakob Prange, Werner Scholze-Stubenrecht, Diana Steffen, Stefan Thater, Christian Wellner, Manfred Pinkal |
| 2016 | LREC | Improving POS Tagging of German Learner Language in a Reading Comprehension Scenario. | Lena Keiper, Andrea Horbach, Stefan Thater |
| 2014 | LREC | Finding a Tradeoff between Accuracy and Rater's Workload in Grading Clustered Short Answers. | Andrea Horbach, Alexis Palmer, Magdalena Wolska |