| 2026 | ACL | Automatic Paper Analysis and Categorisation for Systematic Reviews with Combined Reasoning-Augmented SFT and DAPO RL. | Michela Lorandi, Anya Belz, Simon Mille, Craig Thomson |
| 2026 | ACL | LLM Multi-Agent Systems for Long Triple Set Data-to-Text Generation. | Chinonso Cynthia Osuji, Simon Mille, Mark Andrade, Jane Adkins, Ornait O'Connell, Elaine U Dhonnchadha, Blithn Heffernan, Frinne Nic an tSaoir, Anya Belz, Thiago Castro Ferreira, Brian Davis |
| 2025 | ACL | Standard Quality Criteria Derived from Current NLP Evaluations for Guiding Evaluation Design and Grounding Comparability and AI Compliance Assessments. | Anya Belz, Simon Mille, Craig Thomson |
| 2025 | INLG | Assessing Semantic Consistency in Data-to-Text Generation: A Meta-Evaluation of Textual, Semantic and Model-Based Metrics. | Rudali Huidrom, Michela Lorandi, Simon Mille, Craig Thomson, Anya Belz |
| 2025 | INLG | Scaling Up Data-to-Text Generation to Longer Sequences: A New Dataset and Benchmark Results for Generation from Large Triple Sets. | Chinonso Cynthia Osuji, Simon Mille, Ornait O'Connell, Thiago Castro Ferreira, Anya Belz, Brian Davis |
| 2024 | INLG | QCET: An Interactive Taxonomy of Quality Criteria for Comparable and Repeatable Evaluation of NLP Systems. | Anya Belz, Simon Mille, Craig Thomson, Rudali Huidrom |
| 2024 | INLG | Filling Gaps in Wikipedia: Leveraging Data-to-Text Generation to Improve Encyclopedic Coverage of Underrepresented Groups. | Simon Mille, Massimiliano Pronesti, Craig Thomson, Michela Lorandi, Sophie Fitzpatrick, Rudali Huidrom, Mohammed Sabry, Amy O'Riordan, Anya Belz |
| 2024 | NAACL | On the Role of Summary Content Units in Text Summarization Evaluation. | Marcel Nawrath, Agnieszka Nowak, Tristan Ratz, Danilo C. Walenta, Juri Opitz, Leonardo F. R. Ribeiro, Joo Sedoc, Daniel Deutsch, Simon Mille, Yixin Liu, Sebastian Gehrmann, Lining Zhang, Saad Mahamood, Miruna Clinciu, Khyathi Raghavi Chandu, Yufang Hou |
| 2023 | ACL | Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP. | Anya Belz, Craig Thomson, Ehud Reiter, Simon Mille |
| 2023 | ACL | A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization. | Lining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, Joo Sedoc |
| 2023 | INLG | Mod-D2T: A Multi-layer Dataset for Modular Data-to-Text Generation. | Simon Mille, Franois Lareau, Stamatia Dasiopoulou, Anya Belz |
| 2022 | ACL | Quantified Reproducibility Assessment of NLP Results. | Anya Belz, Maja Popovic, Simon Mille |
| 2021 | ACL | Assessing the Syntactic Capabilities of Transformer-based Multilingual Language Models. | Laura Prez-Mayos, Alba Tboas Garca, Simon Mille, Leo Wanner |
| 2021 | INLG | Text-in-Context: Token-Level Error Detection for Table-to-Text Generation. | Zdenek Kasner, Simon Mille, Ondrej Dusek |
| 2021 | INLG | Another PASS: A Reproduction Study of the Human Evaluation of a Football Report Generation System. | Simon Mille, Thiago Castro Ferreira, Anya Belz, Brian Davis |
| 2020 | INLG | Disentangling the Properties of Human Evaluation Methods: A Classification System to Support Comparability, Meta-Evaluation and Reproducibility Testing. | Anya Belz, Simon Mille, David M. Howcroft |
| 2020 | INLG | Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised Definitions. | David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser |
| 2019 | EMNLP | Back-Translation as Strategy to Tackle the Lack of Corpus in Natural Language Generation from Semantic Representations. | Marco Antonio Sobrevilla Cabezudo, Simon Mille, Thiago A. S. Pardo |
| 2019 | EMNLP | The Second Multilingual Surface Realisation Shared Task (SR'19): Overview and Evaluation Results. | Simon Mille, Anja Belz, Bernd Bohnet, Yvette Graham, Leo Wanner |
| 2019 | INLG | Teaching FORGe to Verbalize DBpedia Properties in Spanish. | Simon Mille, Stamatia Dasiopoulou, Beatrz Fisas, Leo Wanner |
| 2019 | SAC | A portable grammar-based NLG system for verbalization of structured data. | Simon Mille, Stamatia Dasiopoulou, Leo Wanner |
| 2018 | INLG | Underspecified Universal Dependency Structures as Inputs for Multilingual Surface Realisation. | Simon Mille, Anja Belz, Bernd Bohnet, Leo Wanner |
| 2018 | INLG | Sentence Packaging in Text Generation from Semantic Graphs as a Community Detection Problem. | Alexander V. Shvets, Simon Mille, Leo Wanner |
| 2018 | ISMAR | V4Design for Enhancing Architecture and Video Game Creation. | Konstantinos Avgerinakis, Georgios Meditskos, Jens Derdaele, Simon Mille, Yash Shekhawat, Luis Edgardo Fraguada, Eva Lpez, Jolan Wuyts, Anastasios Tellios, Steffen Riegas, Jesper Wachtmeister, Kriszta Doczy, Victor-Jan Vos, Nicolaus Heise, Jens Piesk, Maarten Vergauwen, Leo Wanner, Stefanos Vrochidis, Ioannis Kompatsiaris |
| 2017 | INLG | Shared Task Proposal: Multilingual Surface Realization Using Universal Dependency Trees. | Simon Mille, Bernd Bohnet, Leo Wanner, Anja Belz |
| 2017 | INLG | A demo of FORGe: the Pompeu Fabra Open Rule-based Generator. | Simon Mille, Leo Wanner |
| 2017 | PAAMS | KRISTINA: A Knowledge-Based Virtual Conversation Agent. | Leo Wanner, Elisabeth Andr, Josep Blat, Stamatia Dasiopoulou, Mireia Farrs, Thiago Fraga-Silva, Eleni Kamateri, Florian Lingenfelser, Gerard Llorach, Oriol Martnez, Georgios Meditskos, Simon Mille, Wolfgang Minker, Louisa Pragst, Dominik Schiller, Andries Stam, Ludo Stellingwerff, Federico Sukno, Bianca Vieru, Stefanos Vrochidis |
| 2016 | ECAI | Multilingual Natural Language Generation within Abstractive Summarization. | Simon Mille, Miguel Ballesteros, Alicia Burga, Gerard Casamayor, Leo Wanner |
| 2015 | NAACL | Data-driven sentence generation with non-isomorphic trees. | Miguel Ballesteros, Bernd Bohnet, Simon Mille, Leo Wanner |
| 2015 | NAACL | Visualizing Deep-Syntactic Parser Output. | Juan Soler Company, Miguel Ballesteros, Bernd Bohnet, Simon Mille, Leo Wanner |
| 2014 | COLING | Deep-Syntactic Parsing. | Miguel Ballesteros, Bernd Bohnet, Simon Mille, Leo Wanner |
| 2014 | INLG | Classifiers for data-driven deep sentence generation. | Miguel Ballesteros, Simon Mille, Leo Wanner |
| 2012 | COLING | How Does the Granularity of an Annotation Scheme Influence Dependency Parsing Performance? | Simon Mille, Alicia Burga, Gabriela Ferraro, Leo Wanner |
| 2012 | INLG | The Surface Realisation Task: Recent Developments and Future Plans. | Anja Belz, Bernd Bohnet, Simon Mille, Leo Wanner, Michael White |
| 2012 | INLG | Towards a Surface Realization-Oriented Corpus Annotation. | Leo Wanner, Simon Mille, Bernd Bohnet |
| 2012 | LREC | Text Simplification Tools for Spanish. | Stefan Bott, Horacio Saggion, Simon Mille |
| 2012 | NLDB | From Ontology to NL: Generation of Multilingual User-Oriented Environmental Reports. | Nadjet Bouayad-Agha, Gerard Casamayor, Simon Mille, Marco Rospocher, Horacio Saggion, Luciano Serafini, Leo Wanner |
| 2010 | COLING | Broad Coverage Multilingual Deep Sentence Generation with a Stochastic Multi-Level Realizer. | Bernd Bohnet, Leo Wanner, Simon Mille, Alicia Burga |
| 2010 | LREC | Syntactic Dependencies for Multilingual and Multilevel Corpus Annotation. | Simon Mille, Leo Wanner |
| 2009 | ICAIL | Improving the comprehension of legal documentation: the case of patent claims. | Nadjet Bouayad-Agha, Gerard Casamayor, Gabriela Ferraro, Simon Mille, Vanesa Vidal, Leo Wanner |
| 2008 | EAMT | Multilingual summarization in practice: the case of patent claims. | Simon Mille, Leo Wanner |
| 2008 | LREC | Making Text Resources Accessible to the Reader: the Case of Patent Claims. | Simon Mille, Leo Wanner |