| 2026 | ECIR | The Effect of Document Summarization on LLM-Based Relevance Judgments. | Samaneh Mohtadi, Kevin Roitero, Stefano Mizzaro, Gianluca Demartini |
| 2026 | ECIR | Analyzing AI Evaluation Benchmarks Through Information Retrieval and Network Science. | Gaia Simeoni, Michael Soprano, Riccardo Lunardi, Kevin Roitero, Stefano Mizzaro |
| 2026 | ECIR | Large Language Models as Assessors: On the Impact of Relevance Scales. | Riccardo Zamolo, Riccardo Lunardi, Michael Soprano, Gianluca Demartini, Stefano Mizzaro, Kevin Roitero |
| 2025 | ECAI | On Robustness and Reliability of Benchmark-Based Evaluation of LLMs. | Riccardo Lunardi, Vincenzo Della Mea, Stefano Mizzaro, Kevin Roitero |
| 2025 | ECIR | A Comparative Analysis of Retrieval-Augmented Generation and Crowdsourcing for Fact-Checking. | Francesco Bombassei De Bona, David La Barbera, Stefano Mizzaro, Kevin Roitero |
| 2025 | ECIR | Leveraging LLMs for Energy Forecasting: The AcegasApsAmga Case Study. | Kevin Roitero, Andrea Zancola, Vincenzo Della Mea, Stefano Mizzaro |
| 2025 | ICTIR | Preaching to the ChoIR: Lessons IR Should Share with AI. | Gianluca Demartini, Claudia Hauff, Matthew Lease, Stefano Mizzaro, Kevin Roitero, Mark Sanderson, Falk Scholer, Chirag Shah, Damiano Spina, Paul Thomas, Arjen P. de Vries, Guido Zuccon |
| 2025 | ICTIR | AIDME: A Scalable, Interpretable Framework for AI-Aided Scoping Reviews. | Michael Soprano, Sandip Modha, Kevin Roitero, Eddy Maddalena, Marco Viviani, Gabriella Pasi, Stefano Mizzaro |
| 2025 | SIGIR | PILs of Knowledge: A Synthetic Benchmark for Evaluating Question Answering Systems in Healthcare. | Riccardo Lunardi, Michael Soprano, Paolo Coppola, Vincenzo Della Mea, Stefano Mizzaro, Kevin Roitero |
| 2025 | SIGIR | Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking. | Kevin Roitero, Dustin Wright, Michael Soprano, Isabelle Augenstein, Stefano Mizzaro |
| 2025 | SIGIR | The Magnitude of Truth: On Using Magnitude Estimation for Truthfulness Assessment. | Michael Soprano, Denis Eduard Tapu, David La Barbera, Kevin Roitero, Stefano Mizzaro |
| 2024 | CIKM | Generative AI for Energy: Multi-Horizon Power Consumption Forecasting using Large Language Models. | Kevin Roitero, Gianluca D'Abrosca, Andrea Zancola, Vincenzo Della Mea, Stefano Mizzaro |
| 2024 | SIGIR | Combining Large Language Models and Crowdsourcing for Hybrid Human-AI Misinformation Detection. | Xia Zeng, David La Barbera, Kevin Roitero, Arkaitz Zubiaga, Stefano Mizzaro |
| 2022 | WWW | Preferences on a Budget: Prioritizing Document Pairs when Crowdsourcing Relevance Judgments. | Kevin Roitero, Alessandro Checco, Stefano Mizzaro, Gianluca Demartini |
| 2022 | SIGIR | Ranking Interruptus: When Truncated Rankings Are Better and How to Measure That. | Enrique Amig, Stefano Mizzaro, Damiano Spina |
| 2022 | WSDM | Crowd_Frame: A Simple and Complete Framework to Deploy Complex Crowdsourcing Tasks Off-the-shelf. | Michael Soprano, Kevin Roitero, Francesco Bombassei De Bona, Stefano Mizzaro |
| 2020 | ACL | An Effectiveness Metric for Ordinal Classification: Formal Properties and Experimental Results. | Enrique Amig, Julio Gonzalo, Stefano Mizzaro, Jorge Carrillo-de-Albornoz |
| 2020 | CIKM | The COVID-19 Infodemic: Can the Crowd Judge Recent Misinformation Objectively? | Kevin Roitero, Michael Soprano, Beatrice Portelli, Damiano Spina, Vincenzo Della Mea, Giuseppe Serra, Stefano Mizzaro, Gianluca Demartini |
| 2020 | ECIR | Crowdsourcing Truthfulness: The Impact of Judgment Scale and Assessor Bias. | David La Barbera, Kevin Roitero, Gianluca Demartini, Stefano Mizzaro, Damiano Spina |
| 2020 | ECIR | Twitter goes to the Doctor: Detecting Medical Tweets using Machine Learning and BERT. | Kevin Roitero, Cristian Bozzato, Vincenzo Della Mea, Stefano Mizzaro, Giuseppe Serra |
| 2020 | SIGIR | Can The Crowd Identify Misinformation Objectively?: The Effects of Judgment Scale and Assessor's Background. | Kevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina, Stefano Mizzaro, Gianluca Demartini |
| 2019 | CIKM | On Transforming Relevance Scales. | Lei Han, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Gianluca Demartini |
| 2019 | CIKM | Towards Stochastic Simulations of Relevance Profiles. | Kevin Roitero, Andrea Brunello, Julin Urbano, Stefano Mizzaro |
| 2019 | SIGIR | HITS Hits Readersourcing: Validating Peer Review Alternatives Using Network Analysis. | Michael Soprano, Kevin Roitero, Stefano Mizzaro |
| 2019 | SIGIR | On Topic Difficulty in IR Evaluation: The Effect of Systems, Corpora, and System Components. | Fabio Zampieri, Kevin Roitero, J. Shane Culpepper, Oren Kurland, Stefano Mizzaro |
| 2018 | CIKM | Multidimensional News Quality: A Comparison of Crowdsourcing and Nichesourcing. | Eddy Maddalena, Davide Ceolin, Stefano Mizzaro |
| 2018 | CIKM | How Many Truth Levels? Six? One Hundred? Even More? Validating Truthfulness of Statements via Crowdsourcing. | Kevin Roitero, Gianluca Demartini, Stefano Mizzaro, Damiano Spina |
| 2018 | ICTIR | A Formal Account of Effectiveness Evaluation and Ranking Fusion. | Enrique Amig, Fernando Giner, Stefano Mizzaro, Damiano Spina |
| 2018 | SIGIR | Are we on the Right Track?: An Examination of Information Retrieval Methodologies. | Enrique Amig, Hui Fang, Stefano Mizzaro, ChengXiang Zhai |
| 2018 | SIGIR | Query Performance Prediction and Effectiveness Evaluation Without Relevance Judgments: Two Sides of the Same Coin. | Stefano Mizzaro, Josiane Mothe, Kevin Roitero, Md. Zia Ullah |
| 2018 | SIGIR | On Fine-Grained Relevance Scales. | Kevin Roitero, Eddy Maddalena, Gianluca Demartini, Stefano Mizzaro |
| 2018 | SIGIR | IRevalOO: An Object Oriented Framework for Retrieval Evaluation. | Kevin Roitero, Eddy Maddalena, Yannick Ponte, Stefano Mizzaro |
| 2018 | SIGIR | Effectiveness Evaluation with a Subset of Topics: A Practical Approach. | Kevin Roitero, Michael Soprano, Stefano Mizzaro |
| 2017 | ECIR | Human-Based Query Difficulty Prediction. | Adrian-Gabriel Chifu, Sbastien Djean, Stefano Mizzaro, Josiane Mothe |
| 2017 | ECIR | Do Easy Topics Predict Effectiveness Better Than Difficult Topics? | Kevin Roitero, Eddy Maddalena, Stefano Mizzaro |
| 2017 | HCOMP | Let's Agree to Disagree: Fixing Agreement Measures for Crowdsourcing. | Alessandro Checco, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Gianluca Demartini |
| 2017 | ICTIR | Considering Assessor Agreement in IR Evaluation. | Eddy Maddalena, Kevin Roitero, Gianluca Demartini, Stefano Mizzaro |
| 2017 | SIGIR | Axiomatic Thinking for Information Retrieval: And Related Tasks. | Enrique Amig, Hui Fang, Stefano Mizzaro, ChengXiang Zhai |
| 2016 | ECIR | Exploiting News to Categorize Tweets: Quantifying the Impact of Different News Collections. | Marco Pavan, Stefano Mizzaro, Matteo Bernardon, Ivan Scagnetto |
| 2016 | HCOMP | Crowdsourcing Relevance Assessments: The Unexpected Benefits of Limiting the Time to Judge. | Eddy Maddalena, Marco Basaldella, Dario De Nart, Dante Degl'Innocenti, Stefano Mizzaro, Gianluca Demartini |
| 2016 | SIGIR | Why do you Think this Query is Difficult?: A User Study on Human Query Prediction. | Stefano Mizzaro, Josiane Mothe |
| 2015 | ECIR | A Formal Approach to Effectiveness Metrics for Information Access: Retrieval, Filtering, and Clustering. | Enrique Amig, Julio Gonzalo, Stefano Mizzaro |
| 2015 | ECIR | Different Rankers on Different Subcollections. | Timothy Jones, Falk Scholer, Andrew Turpin, Stefano Mizzaro, Mark Sanderson |
| 2015 | ECIR | Judging Relevance Using Magnitude Estimation. | Eddy Maddalena, Stefano Mizzaro, Falk Scholer, Andrew Turpin |
| 2015 | ECIR | Content-Based Similarity of Twitter Users. | Stefano Mizzaro, Marco Pavan, Ivan Scagnetto |
| 2015 | MDM | Finding Important Locations: A Feature-Based Approach. | Marco Pavan, Stefano Mizzaro, Ivan Scagnetto, Andrea Beggiato |
| 2015 | SIGIR | The Benefits of Magnitude Estimation Relevance Assessments for Information Retrieval Evaluation. | Andrew Turpin, Falk Scholer, Stefano Mizzaro, Eddy Maddalena |
| 2014 | CIKM | Size and Source Matter: Understanding Inconsistencies in Test Collection-Based Evaluation. | Timothy Jones, Andrew Turpin, Stefano Mizzaro, Falk Scholer, Mark Sanderson |
| 2014 | SIGIR | A general account of effectiveness metrics for information tasks: retrieval, filtering, and clustering. | Enrique Amig, Julio Gonzalo, Stefano Mizzaro |
| 2014 | SIGIR | TREC: topic engineering exercise. | J. Shane Culpepper, Stefano Mizzaro, Mark Sanderson, Falk Scholer |
| 2014 | SIGIR | Short text categorization exploiting contextual enrichment and external knowledge. | Stefano Mizzaro, Marco Pavan, Ivan Scagnetto, Martino Valenti |
| 2013 | ICTIR | On Using Fewer Topics in Information Retrieval Evaluations. | Andrea Berto, Stefano Mizzaro, Stephen Robertson |
| 2013 | ICTIR | Axiometrics: An Axiomatic Approach to Information Retrieval Effectiveness Metrics. | Luca Busin, Stefano Mizzaro |
| 2009 | ICTIR | IR Evaluation without a Common Set of Topics. | Matteo Cattelan, Stefano Mizzaro |
| 2009 | ICTIR | Evaluating Mobile Proactive Context-Aware Retrieval: An Incremental Benchmark. | Davide Menegon, Stefano Mizzaro, Elena Nazzi, Luca Vassena |
| 2009 | WWW | QuWi: quality control in Wikipedia. | Alberto Cusinato, Vincenzo Della Mea, Francesco Di Salvatore, Stefano Mizzaro |
| 2009 | SIGIR | Relevance criteria for e-commerce: a crowdsourcing-based experimental analysis. | Omar Alonso, Stefano Mizzaro |
| 2008 | ECAI | AI on the Move: Exploiting AI Techniques for Context Inference on Mobile Devices. | Adolfo Bulfoni, Paolo Coppola, Vincenzo Della Mea, Luca Di Gaspero, Danny Mischis, Stefano Mizzaro, Ivan Scagnetto, Luca Vassena |
| 2008 | ECIR | The Good, the Bad, the Difficult, and the Easy: Something Wrong with Information Retrieval Evaluation?. | Stefano Mizzaro |
| 2008 | WWW | m-Dvara 2.0: Mobile & Web 2.0 Services Integration for Cultural Heritage. | Paolo Coppola, Raffaella Lomuscio, Stefano Mizzaro, Elena Nazzi |
| 2007 | SIGIR | Hits hits TREC: exploring IR evaluation results with network analysis. | Stefano Mizzaro, Stephen Robertson |
| 2006 | ECIR | Mobile Clustering Engine. | Claudio Carpineto, Andrea Della Pietra, Stefano Mizzaro, Giovanni Romano |
| 2006 | ECIR | A Classification of IR Effectiveness Metrics. | Gianluca Demartini, Stefano Mizzaro |
| 2006 | ECIR | Experiments on Average Distance Measure. | Vincenzo Della Mea, Gianluca Demartini, Luca Di Gaspero, Stefano Mizzaro |
| 2000 | EJC | Towards a Theory of Epistemic Information. | Stefano Mizzaro |
| 1996 | SIGIR | Evaluating User Interfaces to Information Retrieval Systems: A Case Study on User Support. | Giorgio Brajnik, Stefano Mizzaro, Carlo Tasso |