| 2026 | ACL | Automatic Paper Analysis and Categorisation for Systematic Reviews with Combined Reasoning-Augmented SFT and DAPO RL. | Michela Lorandi, Anya Belz, Simon Mille, Craig Thomson |
| 2026 | ACL | LLM Multi-Agent Systems for Long Triple Set Data-to-Text Generation. | Chinonso Cynthia Osuji, Simon Mille, Mark Andrade, Jane Adkins, Ornait O'Connell, Elaine U Dhonnchadha, Blithn Heffernan, Frinne Nic an tSaoir, Anya Belz, Thiago Castro Ferreira, Brian Davis |
| 2026 | ACL | Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning. | Massimiliano Pronesti, Anya Belz, Yufang Hou |
| 2026 | ACL | AutoForest: Automatically Generating Forest Plots from Biomedical Studies with End-to-End Evidence Extraction and Synthesis. | Massimiliano Pronesti, Angelo Miculescu, Mohsin Kapdi, Paul Flanagan, Oisin Redmond, Joao H. Bettencourt-Silva, Gurdeep Singh Mannu, Spiros Denaxas, Rui Bebiano Da Providencia E. Costa, Anya Belz, Yufang Hou |
| 2025 | ACL | Standard Quality Criteria Derived from Current NLP Evaluations for Guiding Evaluation Design and Grounding Comparability and AI Compliance Assessments. | Anya Belz, Simon Mille, Craig Thomson |
| 2025 | ACL | Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies. | Massimiliano Pronesti, Joao H. Bettencourt-Silva, Paul Flanagan, Alessandra Pascale, Oisin Redmond, Anya Belz, Yufang Hou |
| 2025 | EMNLP | Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning. | Massimiliano Pronesti, Michela Lorandi, Paul Flanagan, Oisin Redmond, Anya Belz, Yufang Hou |
| 2025 | EMNLP | Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers' Views and Experience of Reproducibility. | Craig Thomson, Ehud Reiter, Joo Sedoc, Anya Belz |
| 2025 | INLG | Assessing Semantic Consistency in Data-to-Text Generation: A Meta-Evaluation of Textual, Semantic and Model-Based Metrics. | Rudali Huidrom, Michela Lorandi, Simon Mille, Craig Thomson, Anya Belz |
| 2025 | INLG | Scaling Up Data-to-Text Generation to Longer Sequences: A New Dataset and Benchmark Results for Generation from Large Triple Sets. | Chinonso Cynthia Osuji, Simon Mille, Ornait O'Connell, Thiago Castro Ferreira, Anya Belz, Brian Davis |
| 2024 | ACL | Beyond Abstracts: A New Dataset, Prompt Design Strategy and Method for Biomedical Synthesis Generation. | James O'Doherty, Cian Nolan, Yufang Hou, Anya Belz |
| 2024 | EACL | High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models. | Michela Lorandi, Anya Belz |
| 2024 | EACL | Assessing the Portability of Parameter Matrices Trained by Parameter-Efficient Finetuning Methods. | Mohammed Mohammed, Anya Belz |
| 2024 | INLG | QCET: An Interactive Taxonomy of Quality Criteria for Comparable and Repeatable Evaluation of NLP Systems. | Anya Belz, Simon Mille, Craig Thomson, Rudali Huidrom |
| 2024 | INLG | Differences in Semantic Errors Made by Different Types of Data-to-text Systems. | Rudali Huidrom, Anya Belz, Michela Lorandi |
| 2024 | INLG | Filling Gaps in Wikipedia: Leveraging Data-to-Text Generation to Improve Encyclopedic Coverage of Underrepresented Groups. | Simon Mille, Massimiliano Pronesti, Craig Thomson, Michela Lorandi, Sophie Fitzpatrick, Rudali Huidrom, Mohammed Sabry, Amy O'Riordan, Anya Belz |
| 2024 | INLG | (Mostly) Automatic Experiment Execution for Human Evaluations of NLP Systems. | Craig Thomson, Anya Belz |
| 2023 | ACL | Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP. | Anya Belz, Craig Thomson, Ehud Reiter, Simon Mille |
| 2023 | ACL | Exploring Variation of Results from Different Experimental Conditions. | Maja Popovic, Mohammad Arvan, Natalie Parde, Anya Belz |
| 2023 | INLG | Mod-D2T: A Multi-layer Dataset for Modular Data-to-Text Generation. | Simon Mille, Franois Lareau, Stamatia Dasiopoulou, Anya Belz |
| 2023 | RANLP | Towards a Consensus Taxonomy for Annotating Errors in Automatically Generated Text. | Rudali Huidrom, Anya Belz |
| 2022 | ACL | Quantified Reproducibility Assessment of NLP Results. | Anya Belz, Maja Popovic, Simon Mille |
| 2022 | ACL | Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation. | Francesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, Aleksandar Savkov |
| 2022 | EMNLP | Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation. | Aleksandar Savkov, Francesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Anya Belz, Ehud Reiter |
| 2022 | NAACL | User-Driven Research of Medical Note Generation Software. | Tom Knoll, Francesco Moramarco, Alex Papadopoulos-Korfiatis, Rachel Young, Claudia Ruffini, Mark Perera, Christian Perstl, Ehud Reiter, Anya Belz, Aleksandar Savkov |
| 2021 | EACL | A Systematic Review of Reproducibility Research in Natural Language Processing. | Anya Belz, Shubham Agarwal, Anastasia Shimorina, Ehud Reiter |
| 2021 | INLG | The ReproGen Shared Task on Reproducibility of Human Evaluations in NLG: Overview and Results. | Anya Belz, Anastasia Shimorina, Shubham Agarwal, Ehud Reiter |
| 2021 | INLG | Another PASS: A Reproduction Study of the Human Evaluation of a Football Report Generation System. | Simon Mille, Thiago Castro Ferreira, Anya Belz, Brian Davis |
| 2021 | INLG | A Reproduction Study of an Annotation-based Human Evaluation of MT Outputs. | Maja Popovic, Anya Belz |
| 2020 | INLG | ReproGen: Proposal for a Shared Task on Reproducibility of Human Evaluations in NLG. | Anya Belz, Shubham Agarwal, Anastasia Shimorina, Ehud Reiter |
| 2020 | INLG | Disentangling the Properties of Human Evaluation Methods: A Classification System to Support Comparability, Meta-Evaluation and Reproducibility Testing. | Anya Belz, Simon Mille, David M. Howcroft |
| 2020 | INLG | Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised Definitions. | David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser |