| 2025 | INLG | Assessing Semantic Consistency in Data-to-Text Generation: A Meta-Evaluation of Textual, Semantic and Model-Based Metrics. | Rudali Huidrom, Michela Lorandi, Simon Mille, Craig Thomson, Anya Belz |
| 2025 | INLG | Do My Eyes Deceive Me? A Survey of Human Evaluations of Hallucinations in NLG. | Patrcia Schmidtov, Eduardo Cal, Simone Balloccu, Dimitra Gkatzia, Rudali Huidrom, Mateusz Lango, Fahime Same, Vilm Zouhar, Saad Mahamood, Ondrej Dusek |
| 2024 | INLG | QCET: An Interactive Taxonomy of Quality Criteria for Comparable and Repeatable Evaluation of NLP Systems. | Anya Belz, Simon Mille, Craig Thomson, Rudali Huidrom |
| 2024 | INLG | Differences in Semantic Errors Made by Different Types of Data-to-text Systems. | Rudali Huidrom, Anya Belz, Michela Lorandi |
| 2024 | INLG | Filling Gaps in Wikipedia: Leveraging Data-to-Text Generation to Improve Encyclopedic Coverage of Underrepresented Groups. | Simon Mille, Massimiliano Pronesti, Craig Thomson, Michela Lorandi, Sophie Fitzpatrick, Rudali Huidrom, Mohammed Sabry, Amy O'Riordan, Anya Belz |
| 2023 | RANLP | Towards a Consensus Taxonomy for Annotating Errors in Automatically Generated Text. | Rudali Huidrom, Anya Belz |