| 2026 | ACL | GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark. | Lotta Kiefer, Christoph Leiter, Sotaro Takeshita, Elena Schmidt, Steffen Eger |
| 2026 | ACL | Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation. | Ran Zhang, Steffen Eger, Arda Tezcan, Wei Zhao, Simone Paolo Ponzetto, Lieve Macken |
| 2026 | EACL | Emotionally Charged, Logically Blurred: AI-driven Emotional Framing Impairs Human Fallacy Detection. | Yanran Chen, Lynn Greschner, Roman Klinger, Michael Klenk, Steffen Eger |
| 2025 | ACL | Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks. | Yanran Chen, Steffen Eger |
| 2025 | EMNLP | Argument Summarization and its Evaluation in the Era of Large Language Models. | Moritz Altemeyer, Steffen Eger, Johannes Daxenberger, Yanran Chen, Tim Altendorf, Philipp Cimiano, Benjamin Schiller |
| 2025 | EMNLP | Graph-Guided Textual Explanation Generation Framework. | Shuzhou Yuan, Jingyi Sun, Ran Zhang, Michael Frber, Steffen Eger, Pepa Atanasova, Isabelle Augenstein |
| 2025 | EMNLP | LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answering. | Ran Zhang, Wei Zhao, Lieve Macken, Steffen Eger |
| 2025 | ICCV | Tikzero: Zero-Shot Text-Guided Graphics Program Synthesis. | Jonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka, Masao Utiyama, Raj Dabre, Steffen Eger, Simone Paolo Ponzetto |
| 2025 | ICLR | ScImage: How good are multimodal large language models at scientific text-to-image generation? | Leixin Zhang, Steffen Eger, Yinjie Cheng, Weihe Zhai, Jonas Belouadi, Fahimeh Moafian, Zhixue Zhao |
| 2025 | IJCNLP | ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation. | Xiao Wang, Daniil Larionov, Siwei Wu, Yiqi Liu, Steffen Eger, Nafise Sadat Moosavi, Chenghua Lin |
| 2025 | NAACL | PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics. | Daniil Larionov, Steffen Eger |
| 2025 | NAACL | How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs. | Ran Zhang, Wei Zhao, Steffen Eger |
| 2024 | COLING | Dependencies over Times and Tools (DoTT). | Andy Lcking, Giuseppe Abrami, Leon Hammerla, Marc Rahn, Daniel Baumartz, Steffen Eger, Alexander Mehler |
| 2024 | EACL | BMX: Boosting Natural Language Generation Metrics with Explainability. | Christoph Leiter, Hoa Nguyen, Steffen Eger |
| 2024 | EMNLP | Evaluating Diversity in Automatic Poetry Generation. | Yanran Chen, Hannes Grner, Sina Zarrie, Steffen Eger |
| 2024 | EMNLP | Fine-Grained Detection of Solidarity for Women and Migrants in 155 Years of German Parliamentary Debates. | Aida Kostikova, Dominik Beese, Benjamin Paassen, Ole Ptz, Gregor Wiedemann, Steffen Eger |
| 2024 | EMNLP | xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics. | Daniil Larionov, Mikhail Seleznyov, Vasiliy Viskov, Alexander Panchenko, Steffen Eger |
| 2024 | EMNLP | PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation. | Christoph Leiter, Steffen Eger |
| 2024 | ICLR | AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ. | Jonas Belouadi, Anne Lauscher, Steffen Eger |
| 2023 | ACL | ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models. | Jonas Belouadi, Steffen Eger |
| 2023 | ACL | Trade-Offs Between Fairness and Privacy in Language Modeling. | Cleo Matzken, Steffen Eger, Ivan Habernal |
| 2023 | EACL | UScore: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation. | Jonas Belouadi, Steffen Eger |
| 2023 | EACL | DiscoScore: Evaluating Text Generation with BERT and Discourse Coherence. | Wei Zhao, Michael Strube, Steffen Eger |
| 2023 | EMNLP | EffEval: A Comprehensive Evaluation of Efficiency for MT Evaluation Metrics. | Daniil Larionov, Jens Grnwald, Christoph Leiter, Steffen Eger |
| 2022 | ACML | Constrained Density Matching and Modeling for Cross-lingual Alignment of Contextualized Representations. | Wei Zhao, Steffen Eger |
| 2022 | COLING | Layer or Representation Space: What Makes BERT-based Evaluation Metrics Robust? | Doan Nam Long Vu, Nafise Sadat Moosavi, Steffen Eger |
| 2022 | EMNLP | Reproducibility Issues for BERT-based Evaluation Metrics. | Yanran Chen, Jonas Belouadi, Steffen Eger |
| 2021 | ACL | Changes in European Solidarity Before and During COVID-19: Evidence from a Large Crowd- and Expert-Annotated Twitter Dataset. | Alexandra Ils, Dan Liu, Daniela Grunow, Steffen Eger |
| 2021 | ACL | BERT-Defense: A Probabilistic Model Based on BERT to Combat Cognitively Inspired Orthographic Adversarial Attacks. | Yannik Keller, Jan Mackensen, Steffen Eger |
| 2021 | ACL | Better than Average: Paired Evaluation of NLP systems. | Maxime Peyrard, Wei Zhao, Steffen Eger, Robert West |
| 2021 | EMNLP | Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors. | Marvin Kaster, Wei Zhao, Steffen Eger |
| 2021 | INLG | TUDA-Reproducibility @ ReproGen: Replicability of Human Evaluation of Text-to-Text and Concept-to-Text Generation. | Christian Richter, Yanran Chen, Steffen Eger |
| 2020 | ACL | SUPERT: Towards New Frontiers in Unsupervised Evaluation Metrics for Multi-Document Summarization. | Yang Gao, Wei Zhao, Steffen Eger |
| 2020 | ACL | On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation Evaluation. | Wei Zhao, Goran Glavas, Maxime Peyrard, Yang Gao, Robert West, Steffen Eger |
| 2020 | COLING | Vec2Sent: Probing Sentence Embeddings with Natural Language Generation. | Martin Kerscher, Steffen Eger |
| 2020 | COLING | Probing Multilingual BERT for Genetic and Typological Signals. | Taraka Rama, Lisa Beinborn, Steffen Eger |
| 2020 | CoNLL | How to Probe Sentence Embeddings in Low-Resource Languages: On Structural Design Choices for Probing Task Evaluation. | Steffen Eger, Johannes Daxenberger, Iryna Gurevych |
| 2020 | IJCNLP | From Hero to Zroe: A Benchmark of Low-Level Adversarial Attacks. | Steffen Eger, Yannik Benz |
| 2020 | LREC | PO-EMO: Conceptualization, Annotation, and Modeling of Aesthetic Emotions in German and English Poetry. | Thomas N. Haider, Steffen Eger, Evgeny Kim, Roman Klinger, Winfried Menninghaus |
| 2020 | SIGMOD | DBPal: A Fully Pluggable NL2SQL Training Pipeline. | Nathaniel Weir, Prasetya Ajie Utama, Alex Galakatos, Andrew Crotty, Amir Ilkhechi, Shekar Ramaswamy, Rohin Bhushan, Nadja Geisler, Benjamin Httasch, Steffen Eger, Ugur etintemel, Carsten Binnig |
| 2019 | ACL | Towards Scalable and Reliable Capsule Networks for Challenging NLP Applications. | Wei Zhao, Haiyun Peng, Steffen Eger, Erik Cambria, Min Yang |
| 2019 | EMNLP | MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance. | Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, Steffen Eger |
| 2019 | NAACL | Does My Rebuttal Matter? Insights from a Major NLP Conference. | Yang Gao, Steffen Eger, Ilia Kuznetsov, Iryna Gurevych, Yusuke Miyao |
| 2019 | NAACL | Text Processing Like Humans Do: Visually Attacking and Shielding NLP Systems. | Steffen Eger, Gzde Gl Sahin, Andreas Rckl, Ji-Ung Lee, Claudia Schulz, Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, Iryna Gurevych |
| 2018 | COLING | Killing Four Birds with Two Stones: Multi-Task Learning for Non-Literal Language Detection. | Erik-Ln Do Dinh, Steffen Eger, Iryna Gurevych |
| 2018 | COLING | Cross-lingual Argumentation Mining: Machine Translation (and a bit of Projection) is All You Need! | Steffen Eger, Johannes Daxenberger, Christian Stab, Iryna Gurevych |
| 2018 | EMNLP | Is it Time to Swish? Comparing Deep Learning Activation Functions Across NLP tasks. | Steffen Eger, Paul Youssef, Iryna Gurevych |
| 2018 | NAACL | Multi-Task Learning for Argumentation Mining in Low-Resource Settings. | Claudia Schulz, Steffen Eger, Johannes Daxenberger, Tobias Kahse, Iryna Gurevych |
| 2018 | NAACL | ArgumenText: Searching for Arguments in Heterogeneous Sources. | Christian Stab, Johannes Daxenberger, Chris Stahlhut, Tristan Miller, Benjamin Schiller, Christopher Tauchmann, Steffen Eger, Iryna Gurevych |
| 2017 | ACL | Neural End-to-End Learning for Computational Argumentation Mining. | Steffen Eger, Johannes Daxenberger, Iryna Gurevych |
| 2017 | EMNLP | What is the Essence of a Claim? Cross-Domain Claim Identification. | Johannes Daxenberger, Steffen Eger, Ivan Habernal, Christian Stab, Iryna Gurevych |
| 2016 | ACL | On the Linearity of Semantic Change: Investigating Meaning Variation via Dynamic Graph Models. | Steffen Eger, Alexander Mehler |
| 2016 | COLING | Language classification from bilingual word embedding graphs. | Steffen Eger, Armin Hoenen, Alexander Mehler |
| 2016 | COLING | Still not there? Comparing Traditional Sequence-to-Sequence Models to Encoder-Decoder Neural Networks on Monotone String Translation Tasks. | Carsten Schnober, Steffen Eger, Erik-Ln Do Dinh, Iryna Gurevych |
| 2016 | LREC | Lemmatization and Morphological Tagging in German and Latin: A Comparison and a Survey of the State-of-the-art. | Steffen Eger, Rdiger Gleim, Alexander Mehler |
| 2015 | ACL | Multiple Many-to-Many Sequence Alignment for Combining String-Valued Variables: A G2P Experiment. | Steffen Eger |
| 2015 | EMNLP | Do we need bigram alignment models? On the effect of alignment quality on transduction accuracy in G2P. | Steffen Eger |
| 2015 | ICMLA | Complex Decomposition of the Negative Distance Kernel. | Tim vor der Brck, Steffen Eger, Alexander Mehler |
| 2015 | Interspeech | Improving G2p from wiktionary and other (web) resources. | Steffen Eger |
| 2015 | KI | Deriving a Primal Form for the Quadratic Power Kernel. | Tim vor der Brck, Steffen Eger |
| 2012 | COLING | S-Restricted Monotone Alignments: Algorithm, Search Space, and Applications. | Steffen Eger |