| 2025 | KDD | ForTune: Running Offline Scenarios to Estimate Impact on Business Metrics. | Georges Dupret, Konstantin Sozinov, Carmen Barcena Gonzalez, Ziggy Zacks, Amber Yuan, Ben Carterette, Manuel Mai, Andrey Gatash, Gwo Liang Lien, Shubham Bansal, Roberto Sanchis-Ojeda, Mounia Lalmas |
| 2024 | CIKM | PODTILE: Facilitating Podcast Episode Browsing with Auto-generated Chapters. | Azin Ghazimatin, Ekaterina Garmash, Gustavo Penha, Kristen Sheets, Martin Achenbach, Oguz Semerci, Remi Galvez, Marcus Tannenberg, Sahitya Mantravadi, Divya Narayanan, Ofeliya Kalaydzhyan, Douglas Cole, Ben Carterette, Ann Clifton, Paul N. Bennett, Claudia Hauff, Mounia Lalmas |
| 2024 | WWW | Long-term Off-Policy Evaluation and Learning. | Yuta Saito, Himan Abdollahpouri, Jesse Anderton, Ben Carterette, Mounia Lalmas |
| 2023 | CIKM | Graph Learning for Exploratory Query Suggestions in an Instant Search System. | Enrico Palumbo, Andreas Damianou, Alice Wang, Alva Liu, Ghazal Fazelnia, Francesco Fabbri, Rui Ferreira, Fabrizio Silvestri, Hugues Bouchard, Claudia Hauff, Mounia Lalmas, Ben Carterette, Praveen Chandar, David Nyhan |
| 2023 | ICTIR | A is for Adele: An Offline Evaluation Metric for Instant Search. | Negar Arabzadeh, Oleksandra Kmet, Ben Carterette, Charles L. A. Clarke, Claudia Hauff, Praveen Chandar |
| 2022 | CHIIR | CHIIR Workshop on Audio Collection Human Interaction (AudioCHI 2022): http: //speechretrievalworkshop.github.io. | Gareth J. F. Jones, Maria Eskevich, Ben Carterette, Joana Correia, Rosie Jones, Jussi Karlgren, Ian Soboroff |
| 2022 | WWW | Using Survival Models to Estimate User Engagement in Online Experiments. | Praveen Chandar, Brian St. Thomas, Lucas Maystre, Vijay Pappu, Roberto Sanchis-Ojeda, Tiffany Wu, Ben Carterette, Mounia Lalmas, Tony Jebara |
| 2021 | WWW | Estimation of Fair Ranking Metrics with Incomplete Judgments. | mer Kirnap, Fernando Diaz, Asia Biega, Michael D. Ekstrand, Ben Carterette, Emine Yilmaz |
| 2021 | SIGIR | Podcast Metadata and Content: Episode Relevance and Attractiveness in Ad Hoc Search. | Ben Carterette, Rosie Jones, Gareth J. F. Jones, Maria Eskevich, Sravana Reddy, Ann Clifton, Yongze Yu, Jussi Karlgren, Ian Soboroff |
| 2021 | SIGIR | Current Challenges and Future Directions in Podcast Information Access. | Rosie Jones, Hamed Zamani, Markus Schedl, Ching-Wei Chen, Sravana Reddy, Ann Clifton, Jussi Karlgren, Helia Hashemi, Aasish Pappu, Zahra Nazari, Longqi Yang, Oguz Semerci, Hugues Bouchard, Ben Carterette |
| 2020 | CIKM | Evaluating Stochastic Rankings with Expected Exposure. | Fernando Diaz, Bhaskar Mitra, Michael D. Ekstrand, Asia J. Biega, Ben Carterette |
| 2020 | COLING | 100, 000 Podcasts: A Spoken English Document Corpus. | Ann Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu, Rezvaneh Rezapour, Hamed R. Bonab, Maria Eskevich, Gareth J. F. Jones, Jussi Karlgren, Ben Carterette, Rosie Jones |
| 2020 | KDD | Advances in Recommender Systems: From Multi-stakeholder Marketplaces to Automated RecSys. | Rishabh Mehrotra, Ben Carterette, Yong Li, Quanming Yao, Chen Gao, James T. Kwok, Qiang Yang, Isabelle Guyon |
| 2020 | WWW | Leveraging Behavioral Heterogeneity Across Markets for Cross-Market Training of Recommender Systems. | Kevin Roitero, Ben Carterette, Rishabh Mehrotra, Mounia Lalmas |
| 2020 | SIGIR | Bayesian Inferential Risk Evaluation On Multiple IR Systems. | Rodger Benham, Ben Carterette, J. Shane Culpepper, Alistair Moffat |
| 2020 | SIGIR | Recommending Podcasts for Cold-Start Users Based on Music Listening and Taste. | Zahra Nazari, Christophe Charbuillet, Johan Pages, Martin Laurent, Denis Charrier, Briana Vecchione, Ben Carterette |
| 2019 | ADCS | Taking Risks with Confidence. | Rodger Benham, Ben Carterette, Alistair Moffat, J. Shane Culpepper |
| 2019 | ICTIR | Statistical Significance Testing in Theory and in Practice. | Ben Carterette |
| 2019 | ICTIR | From a User Model for Query Sessions to Session Rank Biased Precision (sRBP). | Aldo Lipani, Ben Carterette, Emine Yilmaz |
| 2019 | WSDM | Offline Evaluation to Make Decisions About PlaylistRecommendation Algorithms. | Alois Gruson, Praveen Chandar, Christophe Charbuillet, James McInerney, Samantha Hansen, Damien Tardieu, Ben Carterette |
| 2018 | CIBCB | Improving medical search tasks using learning to rank. | Mohammad Alsulmi, Ben Carterette |
| 2018 | CIKM | Estimating Clickthrough Bias in the Cascade Model. | Praveen Chandar, Ben Carterette |
| 2018 | RecSys | Mixed methods for evaluating user satisfaction. | Jean Garcia-Gathright, Christine Hosey, Brian St. Thomas, Ben Carterette, Fernando Diaz |
| 2018 | SIGIR | Offline Comparative Evaluation with Incremental, Minimally-Invasive Online Feedback. | Ben Carterette, Praveen Chandar |
| 2017 | CHIIR | User Click Detection in Ideal Sessions. | Mustafa Zengin, Ben Carterette |
| 2017 | SIGIR | But Is It Statistically Significant?: Statistical Significance in IR Research, 1995-2014. | Ben Carterette |
| 2017 | SIGIR | Statistical Significance Testing in Information Retrieval: Theory and Practice. | Ben Carterette |
| 2016 | DEXA | Generating Pseudo Search History Data in the Absence of Real Search History. | Ashraf Bah, Ben Carterette |
| 2016 | ICTIR | PDF: A Probabilistic Data Fusion Framework for Retrieval and Ranking. | Ashraf Bah Rabiou, Ben Carterette |
| 2016 | SIGIR | Evaluating Retrieval over Sessions: The TREC Session Track 2011-2014. | Ben Carterette, Paul D. Clough, Mark M. Hall, Evangelos Kanoulas, Mark Sanderson |
| 2015 | CIKM | Learning User Preferences for Topically Similar Documents. | Mustafa Zengin, Ben Carterette |
| 2015 | ICTIR | Statistical Significance Testing in Information Retrieval: Theory and Practice. | Ben Carterette |
| 2015 | ICTIR | Bayesian Inference for Information Retrieval Evaluation. | Ben Carterette |
| 2015 | ICTIR | Dynamic Test Collections for Retrieval Evaluation. | Ben Carterette, Ashraf Bah, Mustafa Zengin |
| 2015 | SIGIR | Document Comprehensiveness and User Preferences in Novelty Search Tasks. | Ashraf Bah, Praveen Chandar, Ben Carterette |
| 2015 | SIGIR | The Best Published Result is Random: Sequential Testing and its Effect on Reported Effectiveness. | Ben Carterette |
| 2014 | SIGIR | Statistical significance testing in information retrieval: theory and practice. | Ben Carterette |
| 2013 | ICTIR | Statistical Significance Testing in Information Retrieval: Theory and Practice. | Ben Carterette |
| 2013 | SIGIR | Preference based evaluation measures for novelty and diversity. | Praveen Chandar, Ben Carterette |
| 2013 | SIGIR | Document features predicting assessor disagreement. | Praveen Chandar, William Webber, Ben Carterette |
| 2013 | SIGIR | An adaptive evidence weighting method for medical record search. | Dongqing Zhu, Ben Carterette |
| 2012 | CIKM | Incorporating variability in user behavior into systems based evaluation. | Ben Carterette, Evangelos Kanoulas, Emine Yilmaz |
| 2012 | CIKM | Predicting baby feeding method from unstructured electronic health record data. | Ashwani Rao, Kristin Maiden, Ben Carterette, Deb Ehrenthal |
| 2012 | CIKM | Alternative assessor disagreement and retrieval depth. | William Webber, Praveen Chandar, Ben Carterette |
| 2012 | CIKM | Combining multi-level evidence for medical record retrieval. | Dongqing Zhu, Ben Carterette |
| 2012 | SIGIR | Advances on the development of evaluation measures. | Ben Carterette, Evangelos Kanoulas, Emine Yilmaz |
| 2012 | SIGIR | Using preference judgments for novel document retrieval. | Praveen Chandar, Ben Carterette |
| 2012 | SIGIR | Using PageRank to infer user preferences. | Praveen Chandar, Ben Carterette |
| 2011 | CIKM | Simulating simple user behavior for system effectiveness evaluation. | Ben Carterette, Evangelos Kanoulas, Emine Yilmaz |
| 2011 | ECIR | A Methodology for Evaluating Aggregated Search Results. | Jaime Arguello, Fernando Diaz, Jamie Callan, Ben Carterette |
| 2011 | ECIR | Within-Document Term-Based Index Pruning with Statistical Hypothesis Testing. | Sree Lekha Thota, Ben Carterette |
| 2011 | ICTIR | Model-Based Inference about IR Systems. | Ben Carterette |
| 2011 | SIGIR | System effectiveness, user models, and user utility: a conceptual framework for investigation. | Ben Carterette |
| 2011 | SIGIR | Evaluating multi-query sessions. | Evangelos Kanoulas, Ben Carterette, Paul D. Clough, Mark Sanderson |
| 2010 | SIGIR | Reusable test collections through experimental design. | Ben Carterette, Evangelos Kanoulas, Virgiliu Pavlu, Hui Fang |
| 2010 | SIGIR | Low cost evaluation in information retrieval. | Ben Carterette, Evangelos Kanoulas, Emine Yilmaz |
| 2010 | SIGIR | The effect of assessor error on IR system evaluation. | Ben Carterette, Ian Soboroff |
| 2010 | SIGIR | Diversification of search results using webgraphs. | Praveen Chandar, Ben Carterette |
| 2010 | WSDM | Measuring the reusability of test collections. | Ben Carterette, Evgeniy Gabrilovich, Vanja Josifovski, Donald Metzler |
| 2009 | CIKM | Probabilistic models of ranking novel documents for faceted topic retrieval. | Ben Carterette, Praveen Chandar |
| 2009 | ECIR | If I Had a Million Queries. | Ben Carterette, Virgiliu Pavlu, Evangelos Kanoulas, Javed A. Aslam, James Allan |
| 2009 | ICTIR | An Analysis of NP-Completeness in Novelty and Diversity Ranking. | Ben Carterette |
| 2009 | SIGIR | On rank correlation and the distance between rankings. | Ben Carterette |
| 2009 | SIGIR | Agreement among statistical significance tests for information retrieval evaluation at varying sample sizes. | Mark D. Smucker, James Allan, Ben Carterette |
| 2008 | ECIR | Here or There. | Ben Carterette, Paul N. Bennett, David Maxwell Chickering, Susan T. Dumais |
| 2008 | SIGIR | Evaluation measures for preference judgments. | Ben Carterette, Paul N. Bennett |
| 2008 | SIGIR | Evaluation over thousands of queries. | Ben Carterette, Virgiliu Pavlu, Evangelos Kanoulas, Javed A. Aslam, James Allan |
| 2007 | CIKM | Semiautomatic evaluation of retrieval systems using document similarities. | Ben Carterette, James Allan |
| 2007 | CIKM | Hypothesis testing with incomplete relevance judgments. | Ben Carterette, Mark D. Smucker |
| 2007 | CIKM | A comparison of statistical significance tests for information retrieval evaluation. | Mark D. Smucker, James Allan, Ben Carterette |
| 2007 | SIGIR | Robust test collections for retrieval evaluation. | Ben Carterette |
| 2006 | ACL | N Semantic Classes are Harder than Two. | Ben Carterette, Rosie Jones, Wiley Greiner, Cory Barr |
| 2006 | SIGIR | Minimal test collections for retrieval evaluation. | Ben Carterette, James Allan, Ramesh K. Sitaraman |
| 2006 | SIGIR | Learning a ranking from pairwise preferences. | Ben Carterette, Desislava Petkova |
| 2005 | CIKM | Incremental test collections. | Ben Carterette, James Allan |
| 2005 | SIGIR | When will information retrieval be "good enough"? | James Allan, Ben Carterette, Joshua Lewis |