| 2026 | ECIR | Analyzing AI Evaluation Benchmarks Through Information Retrieval and Network Science. | Gaia Simeoni, Michael Soprano, Riccardo Lunardi, Kevin Roitero, Stefano Mizzaro |
| 2026 | ECIR | Large Language Models as Assessors: On the Impact of Relevance Scales. | Riccardo Zamolo, Riccardo Lunardi, Michael Soprano, Gianluca Demartini, Stefano Mizzaro, Kevin Roitero |
| 2025 | ECAI | On Robustness and Reliability of Benchmark-Based Evaluation of LLMs. | Riccardo Lunardi, Vincenzo Della Mea, Stefano Mizzaro, Kevin Roitero |
| 2025 | ICTIR | Impersonating the Crowd: Evaluating LLMs' Ability to Replicate Human Judgment in Misinformation Assessment. | David La Barbera, Riccardo Lunardi, Mengdie Zhuang, Kevin Roitero |
| 2025 | WWW | Mapping and Influencing the Political Ideology of Large Language Models using Synthetic Personas. | Pietro Bernardelle, Leon Frhling, Stefano Civelli, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini |
| 2025 | SIGIR | PILs of Knowledge: A Synthetic Benchmark for Evaluating Question Answering Systems in Healthcare. | Riccardo Lunardi, Michael Soprano, Paolo Coppola, Vincenzo Della Mea, Stefano Mizzaro, Kevin Roitero |
| 2024 | CIKM | The Elusiveness of Detecting Political Bias in Language Models. | Riccardo Lunardi, David La Barbera, Kevin Roitero |