| 2026 | ACL | Automatic Paper Analysis and Categorisation for Systematic Reviews with Combined Reasoning-Augmented SFT and DAPO RL. | Michela Lorandi, Anya Belz, Simon Mille, Craig Thomson |
| 2025 | ACL | Standard Quality Criteria Derived from Current NLP Evaluations for Guiding Evaluation Design and Grounding Comparability and AI Compliance Assessments. | Anya Belz, Simon Mille, Craig Thomson |
| 2025 | EMNLP | Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers' Views and Experience of Reproducibility. | Craig Thomson, Ehud Reiter, Joo Sedoc, Anya Belz |
| 2025 | INLG | Assessing Semantic Consistency in Data-to-Text Generation: A Meta-Evaluation of Textual, Semantic and Model-Based Metrics. | Rudali Huidrom, Michela Lorandi, Simon Mille, Craig Thomson, Anya Belz |
| 2024 | INLG | QCET: An Interactive Taxonomy of Quality Criteria for Comparable and Repeatable Evaluation of NLP Systems. | Anya Belz, Simon Mille, Craig Thomson, Rudali Huidrom |
| 2024 | INLG | Filling Gaps in Wikipedia: Leveraging Data-to-Text Generation to Improve Encyclopedic Coverage of Underrepresented Groups. | Simon Mille, Massimiliano Pronesti, Craig Thomson, Michela Lorandi, Sophie Fitzpatrick, Rudali Huidrom, Mohammed Sabry, Amy O'Riordan, Anya Belz |
| 2024 | INLG | (Mostly) Automatic Experiment Execution for Human Evaluations of NLP Systems. | Craig Thomson, Anya Belz |
| 2023 | ACL | Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP. | Anya Belz, Craig Thomson, Ehud Reiter, Simon Mille |
| 2023 | INLG | Enhancing factualness and controllability of Data-to-Text Generation via data Views and constraints. | Craig Thomson, Clment Rebuffel, Ehud Reiter, Laure Soulier, Somayajulu Sripada, Patrick Gallinari |
| 2021 | INLG | Underreporting of errors in NLG output, and what to do about it. | Emiel van Miltenburg, Miruna-Adriana Clinciu, Ondrej Dusek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppnen, Saad Mahamood, Emma Manning, Stephanie Schoch, Craig Thomson, Luou Wen |
| 2021 | INLG | Generation Challenges: Results of the Accuracy Evaluation Shared Task. | Craig Thomson, Ehud Reiter |
| 2021 | SIN | Min-max Training: Adversarially Robust Learning Models for Network Intrusion Detection Systems. | Sam Grierson, Craig Thomson, Pavlos Papadopoulos, Bill Buchanan |
| 2020 | INLG | Shared Task on Evaluating Accuracy. | Ehud Reiter, Craig Thomson |
| 2020 | INLG | A Gold Standard Methodology for Evaluating Accuracy in Data-To-Text Systems. | Craig Thomson, Ehud Reiter |
| 2020 | INLG | Studying the Impact of Filling Information Gaps on the Output Quality of Neural Data-to-Text. | Craig Thomson, Zhijie Zhao, Somayajulu Sripada |
| 2018 | HPCC | Performance Investigation of RPL Routing in Pipeline Monitoring WSNs. | Ahmed Yassin Al-Dubai, Isam Wadhaj, Wajeb Gharibi, Craig Thomson |
| 2018 | INLG | Comprehension Driven Document Planning in Natural Language Generation Systems. | Craig Thomson, Ehud Reiter, Somayajulu Sripada |
| 2016 | AICCSA | Performance evaluation of RPL metrics in environments with strained transmission ranges. | Craig Thomson, Isam Wadhaj, Imed Romdhani, Ahmed Yassin Al-Dubai |