| 2025 | COLING | Exploring the Impact of Language Switching on Personality Traits in LLMs. | Jacopo Amidei, Jose Gregorio Ferreira De S, Rubn Nieto Luna, Andreas Kaltenbrunner |
| 2025 | EMNLP | Coherence of Argumentative Dialogue Snippets: A New Method for Large Scale Evaluation with an Application to Inference Anchoring Theory. | Paul Piwek, Jacopo Amidei, Svetlana Stoyanchev |
| 2024 | ICWSM | A Dataset to Assess Microsoft Copilot Answers in the Context of Swiss, Bavarian and Hessian Elections. | Salvatore Romano, Riccardo Angius, Natalie Kerby, Paul Bouchaud, Jacopo Amidei, Andreas Kaltenbrunner |
| 2022 | EMNLP | Opening up Minds with Argumentative Dialogues. | Youmna Farag, Charlotte O. Brand, Jacopo Amidei, Paul Piwek, Tom Stafford, Svetlana Stoyanchev, Andreas Vlachos |
| 2020 | COLING | Identifying Annotator Bias: A new IRT-based method for bias identification. | Jacopo Amidei, Paul Piwek, Alistair Willis |
| 2020 | COLING | Similarity or deeper understanding? Analyzing the TED-Q dataset of evoked questions. | Matthijs Westera, Jacopo Amidei, Laia Mayol |
| 2019 | INLG | Agreement is overrated: A plea for correlation to assess human evaluation reliability. | Jacopo Amidei, Paul Piwek, Alistair Willis |
| 2019 | INLG | The use of rating and Likert scales in Natural Language Generation human evaluation tasks: A review and some recommendations. | Jacopo Amidei, Paul Piwek, Alistair Willis |
| 2018 | COLING | Rethinking the Agreement in Human Evaluation Tasks. | Jacopo Amidei, Paul Piwek, Alistair Willis |
| 2018 | INLG | Evaluation methodologies in Automatic Question Generation 2013-2018. | Jacopo Amidei, Paul Piwek, Alistair Willis |