| 2025 | ACL | CausalGraphBench: a Benchmark for Evaluating Language Models capabilities of Causal Graph discovery. | Nikolay Babakov, Ehud Reiter, Alberto Bugarn Diz |
| 2025 | ACL | SPHERE: An Evaluation Card for Human-AI Systems. | Dora Zhao, Qianou Ma, Xinran Zhao, Chenglei Si, Chenyang Yang, Ryan Louie, Ehud Reiter, Diyi Yang, Tongshuang Wu |
| 2025 | COLING | Scalability of Bayesian Network Structure Elicitation with Large Language Models: a Novel Methodology and Comparative Analysis. | Nikolay Babakov, Ehud Reiter, Alberto Bugarn Diz |
| 2025 | EMNLP | Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers' Views and Experience of Reproducibility. | Craig Thomson, Ehud Reiter, Joo Sedoc, Anya Belz |
| 2025 | INLG | When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition. | Karen Jia-Hui Li, Simone Balloccu, Ondrej Dusek, Ehud Reiter |
| 2025 | INLG | Input Matters: Evaluating Input Structure's Impact on LLM Summaries of Sports Play-by-Play. | Barkavi Sundararajan, Somayajulu Sripada, Ehud Reiter |
| 2024 | EMNLP | Ask the experts: sourcing a high-quality nutrition counseling dataset through Human-AI collaboration. | Simone Balloccu, Ehud Reiter, Karen Li, Rafael Sargsyan, Vivek Kumar, Diego Reforgiato Recupero, Daniele Riboni, Ondrej Dusek |
| 2024 | NAACL | Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTTo. | Barkavi Sundararajan, Somayajulu Sripada, Ehud Reiter |
| 2023 | ACL | Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP. | Anya Belz, Craig Thomson, Ehud Reiter, Simon Mille |
| 2023 | ACL | Are Experts Needed? On Human Evaluation of Counselling Reflection Generation. | Zixiu Wu, Simone Balloccu, Ehud Reiter, Rim Helaoui, Diego Reforgiato Recupero, Daniele Riboni |
| 2023 | ECAI | Evaluation of Human-Understandability of Global Model Explanations Using Decision Tree. | Adarsa Sivaprasad, Ehud Reiter, Nava Tintarev, Nir Oren |
| 2023 | ICASSP | Smart Selection of Useful Insights from Wearables. | Allmin Pradhap Singh Susaiyah, Aki Hrm, Simone Balloccu, Ehud Reiter, Milan Petkovic |
| 2023 | INLG | Enhancing factualness and controllability of Data-to-Text Generation via data Views and constraints. | Craig Thomson, Clment Rebuffel, Ehud Reiter, Laure Soulier, Somayajulu Sripada, Patrick Gallinari |
| 2022 | ACL | Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation. | Francesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, Aleksandar Savkov |
| 2022 | EMNLP | Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation. | Aleksandar Savkov, Francesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Anya Belz, Ehud Reiter |
| 2022 | ICASSP | Anno-MI: A Dataset of Expert-Annotated Counselling Dialogues. | Zixiu Wu, Simone Balloccu, Vivek Kumar, Rim Helaoui, Ehud Reiter, Diego Reforgiato Recupero, Daniele Riboni |
| 2022 | NAACL | User-Driven Research of Medical Note Generation Software. | Tom Knoll, Francesco Moramarco, Alex Papadopoulos-Korfiatis, Rachel Young, Claudia Ruffini, Mark Perera, Christian Perstl, Ehud Reiter, Anya Belz, Aleksandar Savkov |
| 2021 | EACL | A Systematic Review of Reproducibility Research in Natural Language Processing. | Anya Belz, Shubham Agarwal, Anastasia Shimorina, Ehud Reiter |
| 2021 | INLG | The ReproGen Shared Task on Reproducibility of Human Evaluations in NLG: Overview and Results. | Anya Belz, Anastasia Shimorina, Shubham Agarwal, Ehud Reiter |
| 2021 | INLG | Explaining Decision-Tree Predictions by Addressing Potential Conflicts between Predictions and Plausible Expectations. | Sameen Maruf, Ingrid Zukerman, Ehud Reiter, Gholamreza Haffari |
| 2021 | INLG | Generation Challenges: Results of the Accuracy Evaluation Shared Task. | Craig Thomson, Ehud Reiter |
| 2020 | INLG | Arabic NLG Language Functions. | Wael Abed, Ehud Reiter |
| 2020 | INLG | ReproGen: Proposal for a Shared Task on Reproducibility of Human Evaluations in NLG. | Anya Belz, Shubham Agarwal, Anastasia Shimorina, Ehud Reiter |
| 2020 | INLG | Shared Task on Evaluating Accuracy. | Ehud Reiter, Craig Thomson |
| 2020 | INLG | A Gold Standard Methodology for Evaluating Accuracy in Data-To-Text Systems. | Craig Thomson, Ehud Reiter |
| 2020 | IUI | A NLG Framework for User Tailoring and Profiling in Healthcare. | Simone Balloccu, Steffen Pauws, Ehud Reiter |
| 2020 | IUI | Towards a Generalised Framework for Behaviour Insight Mining (short paper). | Allmin Pradhap Singh Susaiyah, Aki Hrm, Ehud Reiter, Rim Helaoui, Milan Petkovic |
| 2018 | INLG | Generating Summaries of Sets of Consumer Products: Learning from Experiments. | Kittipitch Kuptavanich, Ehud Reiter, Kees van Deemter, Advaith Siddharthan |
| 2018 | INLG | Meteorologists and Students: A resource for language grounding of geographical descriptors. | Alejandro Ramos-Soto, Ehud Reiter, Kees van Deemter, Jos Maria Alonso, Albert Gatt |
| 2018 | INLG | Comprehension Driven Document Planning in Natural Language Generation Systems. | Craig Thomson, Ehud Reiter, Somayajulu Sripada |
| 2017 | INLG | Textually Summarising Incomplete Data. | Stephanie Inglis, Ehud Reiter, Somayajulu Sripada |
| 2017 | INLG | A Commercial Perspective on Reference. | Ehud Reiter |
| 2016 | INLG | Absolute and Relative Properties in Geographic Referring Expressions. | Rodrigo de Oliveira, Somayajulu Sripada, Ehud Reiter |
| 2014 | INLG | Generating Annotated Graphs using the NLG Pipeline Architecture. | Saad Mahamood, William Bradshaw, Ehud Reiter |
| 2013 | CogSci | Typicality and Object Reference. | Margaret Mitchell, Ehud Reiter, Kees van Deemter |
| 2013 | NAACL | Generating Expressions that Refer to Visible Objects. | Margaret Mitchell, Kees van Deemter, Ehud Reiter |
| 2012 | INLG | Working with Clinicians to Improve a Patient-Information NLG System. | Saad Mahamood, Ehud Reiter |
| 2011 | ASSETS | A mobile phone based personal narrative system. | Rolf Black, Annalu Waller, Nava Tintarev, Ehud Reiter, Joseph Reddington |
| 2011 | CogSci | On the Use of Size Modifiers When Referring to Visible Objects. | Margaret Mitchell, Kees van Deemter, Ehud Reiter |
| 2010 | EACL | Generating Approximate Geographic Descriptions. | Ross Turner, Somayajulu Sripada, Ehud Reiter |
| 2010 | INLG | Natural Reference to Objects in a Visual Domain. | Margaret Mitchell, Kees van Deemter, Ehud Reiter |
| 2009 | CHI | Facilitating benign deceit in mediated communication. | Wendy Moncur, Judith Masthoff, Ehud Reiter |
| 2008 | AMIA | Summarising Complex ICU Data in Natural Language. | Jim Hunter, Yvonne Freer, Albert Gatt, Robert Logie, Neil McIntosh, Marian van der Meulen, Franois Portet, Ehud Reiter, Somayajulu Sripada, Cindy Sykes |
| 2008 | CBMS | Neonatal Intensive Care Information for Parents - An Affective Approach. | Saad Mahamood, Ehud Reiter, Chris Mellish |
| 2008 | CBMS | What Do You Want to Know? Investigating the Information Requirements of Patient Supporters. | Wendy Moncur, Judith Masthoff, Ehud Reiter |
| 2008 | ECAI | Using Natural Language Generation Technology to Improve Information Flows in Intensive Care Units. | Jim Hunter, Albert Gatt, Franois Portet, Ehud Reiter, Somayajulu Sripada |
| 2008 | INLG | The Importance of Narrative and Other Lessons from an Evaluation of an NLG System that Summarises Clinical Data. | Ehud Reiter, Albert Gatt, Franois Portet, Marian van der Meulen |
| 2008 | INLG | Using Spatial Reference Frames to Generate Grounded Textual Summaries of Georeferenced Data. | Ross Turner, Somayajulu Sripada, Ehud Reiter, Ian Davy |
| 2007 | AIME | Automatic Generation of Textual Summaries from Neonatal Intensive Care Data. | Franois Portet, Ehud Reiter, Jim Hunter, Somayajulu Sripada |
| 2007 | SGAI | Selecting the Content of Textual Descriptions of Geographically Located Events in Spatio-Temporal Weather Data. | Ross Turner, Somayajulu Sripada, Ehud Reiter, Ian P. Davy |
| 2006 | EACL | Comparing Automatic and Human Evaluation of NLG Systems. | Anja Belz, Ehud Reiter |
| 2006 | EACL | Generating Spatio-Temporal Descriptions in Pollen Forecasts. | Ross Turner, Somayajulu Sripada, Ehud Reiter, Ian P. Davy |
| 2006 | INLG | GENEVAL: A Proposal for Shared-task Evaluation in NLG. | Ehud Reiter, Anja Belz |
| 2005 | IJCAI | Evaluating an NLG System using Post-Editing. | Somayajulu Sripada, Ehud Reiter, Lezan Hawizy |
| 2005 | IJCAI | Appropriate Microplanning Choices for Low-Skilled Readers. | Sandra Williams, Ehud Reiter |
| 2005 | SGAI | Generating Feedback Reports for Adults Taking Basic Skills Tests. | Ehud Reiter, Sandra Williams, Lesley Crichton |
| 2004 | ECAI | Lessons from Deploying NLG Technology for Marine Weather Forecast Text Generation. | Somayajulu Sripada, Ehud Reiter, Ian Davy, Kristian Nilssen |
| 2004 | INLG | Contextual Influences on Near-Synonym Choice. | Ehud Reiter, Somayajulu Sripada |
| 2003 | EACL | Summarizing Neonatal Time Series Data. | Somayajulu Sripada, Ehud Reiter, Jim Hunter, Jin Yu |
| 2003 | KDD | Generating English summaries of time series data using the Gricean maxims. | Somayajulu Sripada, Ehud Reiter, Jim Hunter, Jin Yu |
| 2002 | INLG | Should Corpora Texts Be Gold Standards for NLG? | Ehud Reiter, Somayajulu Sripada |
| 2001 | ACL | Using a Randomised Controlled Clinical Trial to Evaluate an NLG System. | Ehud Reiter, Roma Robertson, A. Scott Lennox, Liesl Osman |
| 2000 | INLG | Knowledge Acquisition for Natural Language Generation. | Ehud Reiter, Roma Robertson, Liesl Osman |
| 1996 | INLG | The ModelExplainer. | Benoit Lavoie, Owen Rambow, Ehud Reiter |
| 1994 | INLG | Has a Consensus NL Generation Architecture Appeared, and is it Psycholinguistically Plausible? | Ehud Reiter |
| 1993 | IJCAI | Using Classification as a Programming Language. | Chris Mellish, Ehud Reiter |
| 1993 | IJCAI | Optimizing the Costs and Benefits of Natural Language Generation. | Ehud Reiter, Chris Mellish |
| 1992 | ACL | Using Classification to Generate Text. | Ehud Reiter, Chris Mellish |
| 1992 | COLING | A Fast Algorithm for the Generation of Referring Expressions. | Ehud Reiter, Robert Dale |
| 1990 | AAAI | Avoiding Unwanted Conversational Implicatures in Text and Graphics. | Joseph Marks, Ehud Reiter |
| 1990 | ACL | The Computational Complexity of Avoiding Conversational Implicatures. | Ehud Reiter |
| 1990 | INLG | A New Model for Lexical Choice for Open-Class Words. | Ehud Reiter |