| 2026 | ACL | Towards Self-Improving Error Diagnosis in Multi-Agent Systems. | Jiazheng Li, Emine Yilmaz, Bei Chen, Thu Le |
| 2026 | ACL | Beyond Static Toolsets: Self-Evolving LLM Tool Agents via Continual Documentation Adaptation. | Bin Wu, Edgar Meij, Emine Yilmaz |
| 2026 | ACL | Mitigating Context Interference for Reliable and Efficient Search Agents. | Boyang Xue, Bin Wu, Shuofei Qiao, Sheng Wang, Rui Wang, Yiming Du, Hongru Wang, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani |
| 2026 | SIGIR | AgentSearch: Indexing, Retrieval, and Ranking of AI Agents. | Bin Wu, To Eun Kim, Yue Feng, Fernando Diaz, Zhaochun Ren, Emine Yilmaz |
| 2025 | ACL | PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants. | Zheng Zhao, Clara Vania, Subhradeep Kayal, Naila Khan, Shay B. Cohen, Emine Yilmaz |
| 2025 | ACL | A Joint Optimization Framework for Enhancing Efficiency of Tool Utilization in LLM Agents. | Bin Wu, Edgar Meij, Emine Yilmaz |
| 2025 | ACL | EventRAG: Enhancing LLM Generation with Event Knowledge Graphs. | Zairun Yang, Yilin Wang, Zhengyan Shi, Yuan Yao, Lei Liang, Keyan Ding, Emine Yilmaz, Huajun Chen, Qiang Zhang |
| 2025 | CIKM | Empirical Analysis on User Profile in Personalized LLMs. | Bin Wu, Zhengyan Shi, Hossein A. Rahmani, Varsha Ramineni, Emine Yilmaz |
| 2025 | CIKM | ProActLLM: Proactive Conversational Information Seeking with Large Language Models. | Shubham Chatterjee, Xi Wang, Shuo Zhang, Sajad Ebrahimi, Zhaochun Ren, Debasis Ganguly, Gareth J. F. Jones, Emine Yilmaz, Hamed Zamani |
| 2025 | CIKM | Towards Understanding Bias in Synthetic Data for Evaluation. | Hossein A. Rahmani, Varsha Ramineni, Emine Yilmaz, Nick Craswell, Bhaskar Mitra |
| 2025 | ECIR | KEIR @ ECIR 2025: The Second Workshop on Knowledge-Enhanced Information Retrieval. | Zihan Wang, Jinyuan Fang, Giacomo Frisoni, Zhuyun Dai, Zaiqiao Meng, Gianluca Moro, Emine Yilmaz |
| 2025 | EMNLP | Machine-generated text detection prevents language model collapse. | George Drayson, Emine Yilmaz, Vasileios Lampos |
| 2025 | NAACL | Adaptive Retrieval-Augmented Generation for Conversational Systems. | Xi Wang, Procheta Sen, Ruizhe Li, Emine Yilmaz |
| 2025 | WWW | SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval. | Hossein A. Rahmani, Xi Wang, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Paul Thomas |
| 2025 | WWW | JudgeBlender: Ensembling Automatic Relevance Judgments. | Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra |
| 2025 | SIGIR | Bridging the Gap: From Ad-hoc to Proactive Search in Conversations. | Chuan Meng, Francesco Tonolini, Fengran Mo, Nikolaos Aletras, Emine Yilmaz, Gabriella Kazai |
| 2025 | SIGIR | LLM4Eval: Large Language Model for Evaluation in IR. | Clemencia Siro, Hossein A. Rahmani, Mohammad Aliannejadi, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, Emine Yilmaz |
| 2025 | WSDM | LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval. | Hossein A. Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, Emine Yilmaz |
| 2024 | AAAI | A Toolbox for Modelling Engagement with Educational Videos. | Yuxiang Qiu, Karim Djemili, Denis Elezi, Aaneel Shalman Srazali, Mara Prez-Ortiz, Emine Yilmaz, John Shawe-Taylor, Sahan Bulathwela |
| 2024 | AIED | Towards Human-Like Educational Question Generation with Small Language Models. | Fares Fawzi, Sarang Balan, Mutlu Cukurova, Emine Yilmaz, Sahan Bulathwela |
| 2024 | EACL | Clarifying the Path to User Satisfaction: An Investigation into Clarification Usefulness. | Hossein A. Rahmani, Xi Wang, Mohammad Aliannejadi, Mohammadmehdi Naghiaei, Emine Yilmaz |
| 2024 | ECIR | KEIR @ ECIR 2024: The First Workshop on Knowledge-Enhanced Information Retrieval. | Zaiqiao Meng, Shangsong Liang, Xin Xin, Gianluca Moro, Evangelos Kanoulas, Emine Yilmaz |
| 2024 | ECIR | Simulated Task Oriented Dialogues for Developing Versatile Conversational Agents. | Xi Wang, Procheta Sen, Ruizhe Li, Emine Yilmaz |
| 2024 | SIGIR | Synthetic Test Collections for Retrieval Evaluation. | Hossein A. Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, Daniel Campos |
| 2024 | SIGIR | LLM4Eval: Large Language Model for Evaluation in IR. | Hossein A. Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, Emine Yilmaz |
| 2023 | AAAI | Pre-training with Scientific Text Improves Educational Question Generation (Student Abstract). | Hamze Muse, Sahan Bulathwela, Emine Yilmaz |
| 2023 | AAAI | Task2KB: A Public Task-Oriented Knowledge Base. | Procheta Sen, Xi Wang, Ruiqing Xu, Emine Yilmaz |
| 2023 | ACL | Modeling User Satisfaction Dynamics in Dialogue via Hawkes Process. | Fanghua Ye, Zhiyuan Hu, Emine Yilmaz |
| 2023 | ACL | Schema-Guided User Satisfaction Modeling for Task-Oriented Dialogues. | Yue Feng, Yunlong Jiao, Animesh Prasad, Nikolaos Aletras, Emine Yilmaz, Gabriella Kazai |
| 2023 | ACL | A Survey on Asking Clarification Questions Datasets in Conversational Systems. | Hossein A. Rahmani, Xi Wang, Yue Feng, Qiang Zhang, Emine Yilmaz, Aldo Lipani |
| 2023 | ACL | Rethinking Semi-supervised Learning with Language Models. | Zhengxiang Shi, Francesco Tonolini, Nikolaos Aletras, Emine Yilmaz, Gabriella Kazai, Yunlong Jiao |
| 2023 | AIED | Scalable Educational Question Generation with Pre-trained Language Models. | Sahan Bulathwela, Hamze Muse, Emine Yilmaz |
| 2023 | CIKM | On the Reliability of User Feedback for Evaluating the Quality of Conversational Agents. | Jordan Massiah, Emine Yilmaz, Yunlong Jiao, Gabriella Kazai |
| 2023 | EMNLP | Enhancing Conversational Search: Large Language Model-Aided Informative Query Rewriting. | Fanghua Ye, Meng Fang, Shenghui Li, Emine Yilmaz |
| 2023 | EMNLP | Improving Conversational Recommendation Systems via Bias Analysis and Language-Model-Enhanced Data Augmentation. | Xi Wang, Hossein A. Rahmani, Jiqun Liu, Emine Yilmaz |
| 2023 | SIGIR | Query-specific Variable Depth Pooling via Query Performance Prediction. | Debasis Ganguly, Emine Yilmaz |
| 2022 | ACL | ASSIST: Towards Label Noise-Robust Dialogue State Tracking. | Fanghua Ye, Yue Feng, Emine Yilmaz |
| 2022 | ACL | Dynamic Schema Graph Fusion Network for Multi-Domain Dialogue State Tracking. | Yue Feng, Aldo Lipani, Fanghua Ye, Qiang Zhang, Emine Yilmaz |
| 2022 | CHIIR | Watch Less and Uncover More: Could Navigation Tools Help Users Search and Explore Videos? | Mara Prez-Ortiz, Sahan Bulathwela, Claire Dormann, Meghana Verma, Stefan Kreitmayer, Richard Noss, John Shawe-Taylor, Yvonne Rogers, Emine Yilmaz |
| 2022 | CIKM | Workshop on Proactive and Agent-Supported Information Retrieval (PASIR). | Gareth J. F. Jones, Procheta Sen, Debasis Ganguly, Emine Yilmaz |
| 2022 | EDM | Can Population-based Engagement Improve Personalisation? A Novel Dataset and Experiments. | Sahan Bulathwela, Meghana Verma, Mara Prez-Ortiz, Emine Yilmaz, John Shawe-Taylor |
| 2022 | EMNLP | MetaASSIST: Robust Dialogue State Tracking with Meta Learning. | Fanghua Ye, Xi Wang, Jie Huang, Shenghui Li, Samuel Stern, Emine Yilmaz |
| 2022 | ICLR | Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations. | Fangyu Liu, Yunlong Jiao, Jordan Massiah, Emine Yilmaz, Serhii Havrylov |
| 2022 | ICTIR | Evaluating the Cranfield Paradigm for Conversational Search Systems. | Xiao Fu, Emine Yilmaz, Aldo Lipani |
| 2022 | ICTIR | From Search Queries to Conversations in the Design of Information Retrieval and Access Systems. | Emine Yilmaz |
| 2022 | WWW | Similarity-based Multi-Domain Dialogue State Tracking with Copy Mechanisms for Task-based Virtual Personal Assistants. | Jarana Manotumruksa, Jeffrey Dalton, Edgar Meij, Emine Yilmaz |
| 2022 | SIGIR | Fostering Coopetition While Plugging Leaks: The Design and Implementation of the MS MARCO Leaderboards. | Jimmy Lin, Daniel Campos, Nick Craswell, Bhaskar Mitra, Emine Yilmaz |
| 2022 | SIGdial | MultiWOZ 2.4: A Multi-Domain Task-Oriented Dialogue Dataset with Essential Annotation Corrections to Improve State Tracking Evaluation. | Fanghua Ye, Jarana Manotumruksa, Emine Yilmaz |
| 2021 | CSCW | Investigating and Mitigating Biases in Crowdsourced Data. | Danula Hettiachchi, Mark Sanderson, Jorge Gonalves, Simo Hosio, Gabriella Kazai, Matthew Lease, Mike Schaekermann, Emine Yilmaz |
| 2021 | ECIR | Detecting and Forecasting Misinformation via Temporal and Geometric Propagation Patterns. | Qiang Zhang, Jonathan Cook, Emine Yilmaz |
| 2021 | EMNLP | Improving Dialogue State Tracking with Turn-based Loss Function and Sequential Data Augmentation. | Jarana Manotumruksa, Jeff Dalton, Edgar Meij, Emine Yilmaz |
| 2021 | IUI | X5Learn: A Personalised Learning Companion at the Intersection of AI and HCI. | Mara Prez-Ortiz, Claire Dormann, Yvonne Rogers, Sahan Bulathwela, Stefan Kreitmayer, Emine Yilmaz, Richard Noss, John Shawe-Taylor |
| 2021 | WWW | Estimation of Fair Ranking Metrics with Incomplete Judgments. | mer Kirnap, Fernando Diaz, Asia Biega, Michael D. Ekstrand, Ben Carterette, Emine Yilmaz |
| 2021 | WWW | Slot Self-Attentive Dialogue State Tracking. | Fanghua Ye, Jarana Manotumruksa, Qiang Zhang, Shenghui Li, Emine Yilmaz |
| 2021 | WWW | Learning Neural Point Processes with Latent Graphs. | Qiang Zhang, Aldo Lipani, Emine Yilmaz |
| 2021 | SIGIR | MS MARCO: Benchmarking Ranking Models in the Large-Data Regime. | Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Jimmy Lin |
| 2021 | SIGIR | TREC Deep Learning Track: Reusable Test Collections in the Large Data Regime. | Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Ellen M. Voorhees, Ian Soboroff |
| 2021 | SIGIR | Significant Improvements over the State of the Art? A Case Study of the MS MARCO Document Ranking Leaderboard. | Jimmy Lin, Daniel Campos, Nick Craswell, Bhaskar Mitra, Emine Yilmaz |
| 2020 | AAAI | Towards an Integrative Educational Recommender for Lifelong Learners (Student Abstract). | Sahan Bulathwela, Mara Prez-Ortiz, Emine Yilmaz, John Shawe-Taylor |
| 2020 | AAAI | TrueLearn: A Family of Bayesian Algorithms to Match Lifelong Learners to Open Educational Resources. | Sahan Bulathwela, Mara Prez-Ortiz, Emine Yilmaz, John Shawe-Taylor |
| 2020 | CIKM | ORCAS: 20 Million Clicked Query-Document Pairs for Analyzing Search. | Nick Craswell, Daniel Campos, Bhaskar Mitra, Emine Yilmaz, Bodo Billerbeck |
| 2020 | EDM | Predicting Engagement in Video Lectures. | Sahan Bulathwela, Mara Prez-Ortiz, Aldo Lipani, Emine Yilmaz, John Shawe-Taylor |
| 2020 | EMNLP | Unsupervised Few-Bits Semantic Hashing with Implicit Topics Modeling. | Fanghua Ye, Jarana Manotumruksa, Emine Yilmaz |
| 2020 | ICML | Self-Attentive Hawkes Process. | Qiang Zhang, Aldo Lipani, mer Kirnap, Emine Yilmaz |
| 2020 | SIGIR | CrossBERT: A Triplet Neural Architecture for Ranking Entity Properties. | Jarana Manotumruksa, Jeff Dalton, Edgar Meij, Emine Yilmaz |
| 2020 | SIGIR | Sequential-based Adversarial Optimisation for Personalised Top-N Item Recommendation. | Jarana Manotumruksa, Emine Yilmaz |
| 2020 | SIGIR | On the Reliability of Test Collections for Evaluating Systems of Different Types. | Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Daniel Campos |
| 2020 | WSDM | SUM'20: State-based User Modelling. | Sahan Bulathwela, Mara Prez-Ortiz, Rishabh Mehrotra, Davor Orlic, Colin de la Higuera, John Shawe-Taylor, Emine Yilmaz |
| 2019 | AAAI | Towards Task Understanding in Visual Settings. | Sebastin Santy, Wazeer Zulfikar, Rishabh Mehrotra, Emine Yilmaz |
| 2019 | ICTIR | From a User Model for Query Sessions to Session Rank Biased Precision (sRBP). | Aldo Lipani, Ben Carterette, Emine Yilmaz |
| 2019 | WWW | From Stances' Imbalance to Their HierarchicalRepresentation and Detection. | Qiang Zhang, Shangsong Liang, Aldo Lipani, Zhaochun Ren, Emine Yilmaz |
| 2019 | WWW | Reply-Aided Detection of Misinformation via Bayesian Deep Learning. | Qiang Zhang, Aldo Lipani, Shangsong Liang, Emine Yilmaz |
| 2019 | SIGIR | An Analysis of the Change in Discussions on Social Media with Bitcoin Price. | Andrew Burnie, Emine Yilmaz |
| 2018 | CHIIR | Study of Relevance and Effort across Devices. | Manisha Verma, Emine Yilmaz, Nick Craswell |
| 2018 | WWW | Ranking-based Method for News Stance Detection. | Qiang Zhang, Emine Yilmaz, Shangsong Liang |
| 2018 | WSDM | LearnIR: WSDM 2018 Workshop on Learning from User Interactions. | Rishabh Mehrotra, Ahmed Hassan Awadallah, Emine Yilmaz |
| 2017 | AAAI | Collaborative User Clustering for Short Text Streams. | Shangsong Liang, Zhaochun Ren, Emine Yilmaz, Evangelos Kanoulas |
| 2017 | CHIIR | Second Workshop on Supporting Complex Search Tasks. | Nicholas J. Belkin, Toine Bogers, Jaap Kamps, Diane Kelly, Marijn Koolen, Emine Yilmaz |
| 2017 | CHIIR | User Behaviour and Task Characteristics: A Field Study of Daily Information Behaviour. | Jiyin He, Emine Yilmaz |
| 2017 | CHIIR | Current Research in Supporting Complex Search Tasks. | Marijn Koolen, Jaap Kamps, Toine Bogers, Nicholas J. Belkin, Diane Kelly, Emine Yilmaz |
| 2017 | CIKM | Deep Sequential Models for Task Satisfaction Prediction. | Rishabh Mehrotra, Ahmed Hassan Awadallah, Milad Shokouhi, Emine Yilmaz, Imed Zitouni, Ahmed El Kholy, Madian Khabsa |
| 2017 | CIKM | Task Embeddings: Learning Query Embeddings using Task Context. | Rishabh Mehrotra, Emine Yilmaz |
| 2017 | ECIR | Search Costs vs. User Satisfaction on Mobile. | Manisha Verma, Emine Yilmaz |
| 2017 | WWW | Auditing Search Engines for Differential Satisfaction Across Demographics. | Rishabh Mehrotra, Ashton Anderson, Fernando Diaz, Amit Sharma, Hanna M. Wallach, Emine Yilmaz |
| 2017 | WWW | A Concept Language Model for Ad-hoc Retrieval. | Bin Zou, Vasileios Lampos, Shangsong Liang, Zhaochun Ren, Emine Yilmaz, Ingemar J. Cox |
| 2017 | SIGIR | Extracting Hierarchies of Search Tasks & Subtasks via a Bayesian Nonparametric Approach. | Rishabh Mehrotra, Emine Yilmaz |
| 2016 | CHIIR | Characterizing Users' Multi-Tasking Behavior in Web Search. | Rishabh Mehrotra, Prasanta Bhattacharya, Emine Yilmaz |
| 2016 | CHIIR | Category Oriented Task Extraction. | Manisha Verma, Emine Yilmaz |
| 2016 | ECIR | Characterizing Relevance on Mobile and Desktop. | Manisha Verma, Emine Yilmaz |
| 2016 | KDD | Dynamic Clustering of Streaming Short Documents. | Shangsong Liang, Emine Yilmaz, Evangelos Kanoulas |
| 2016 | NAACL | Deconstructing Complex Search Tasks: a Bayesian Nonparametric Approach for Extracting Sub-tasks. | Rishabh Mehrotra, Prasanta Bhattacharya, Emine Yilmaz |
| 2016 | SIGIR | Uncovering Task Based Behavioral Heterogeneities in Online Search Behavior. | Rishabh Mehrotra, Prasanta Bhattacharya, Emine Yilmaz |
| 2016 | SIGIR | Bayesian Performance Comparison of Text Classifiers. | Dell Zhang, Jun Wang, Emine Yilmaz, Xiaoling Wang, Yuxin Zhou |
| 2016 | SIGIR | Explainable User Clustering in Short Text Streams. | Yukun Zhao, Shangsong Liang, Zhaochun Ren, Jun Ma, Emine Yilmaz, Maarten de Rijke |
| 2016 | WSDM | On Obtaining Effort Based Judgements for Information Retrieval. | Manisha Verma, Emine Yilmaz, Nick Craswell |
| 2015 | ICTIR | Terms, Topics & Tasks: Enhanced User Modelling for Better Personalization. | Rishabh Mehrotra, Emine Yilmaz |
| 2015 | WWW | Towards Hierarchies of Search Tasks & Subtasks. | Rishabh Mehrotra, Emine Yilmaz |
| 2015 | SIGIR | IR Evaluation: Modeling User Behavior for Measuring Effectiveness. | Charles L. A. Clarke, Mark D. Smucker, Emine Yilmaz |
| 2015 | SIGIR | IR Evaluation: Designing an End-to-End Offline Evaluation Pipeline. | Jin Young Kim, Emine Yilmaz |
| 2015 | SIGIR | Representative & Informative Query Selection for Learning to Rank using Submodular Functions. | Rishabh Mehrotra, Emine Yilmaz |
| 2015 | SIGIR | Anchoring and Adjustment in Relevance Estimation. | Milad Shokouhi, Ryen White, Emine Yilmaz |
| 2014 | CIKM | Entity Oriented Task Extraction from Query Logs. | Manisha Verma, Emine Yilmaz |
| 2014 | CIKM | Effect of Intent Descriptions on Retrieval Evaluation. | Emine Yilmaz, Evangelos Kanoulas, Nick Craswell |
| 2014 | CIKM | Relevance and Effort: An Analysis of Document Utility. | Emine Yilmaz, Manisha Verma, Nick Craswell, Filip Radlinski, Peter Bailey |
| 2014 | RecSys | Task-Based User Modelling for Personalization via Probabilistic Matrix Factorization. | Rishabh Mehrotra, Emine Yilmaz, Manisha Verma |
| 2013 | CIKM | User intent and assessor disagreement in web search evaluation. | Gabriella Kazai, Emine Yilmaz, Nick Craswell, Seyed M. M. Tahaghoghi |
| 2013 | ICTIR | Modelling Score Distributions Without Actual Scores. | Stephen Robertson, Evangelos Kanoulas, Emine Yilmaz |
| 2013 | SIGIR | SIGIR 2013 workshop on modeling user behavior for information retrieval evaluation. | Charles L. A. Clarke, Luanne Freund, Mark D. Smucker, Emine Yilmaz |
| 2012 | CIKM | Incorporating variability in user behavior into systems based evaluation. | Ben Carterette, Evangelos Kanoulas, Emine Yilmaz |
| 2012 | CIKM | An analysis of systematic judging errors in information retrieval. | Gabriella Kazai, Nick Craswell, Emine Yilmaz, Seyed M. M. Tahaghoghi |
| 2012 | SIGIR | Advances on the development of evaluation measures. | Ben Carterette, Evangelos Kanoulas, Emine Yilmaz |
| 2012 | SIGIR | An uncertainty-aware query selection model for evaluation of IR systems. | Mehdi Hosseini, Ingemar J. Cox, Natasa Milic-Frayling, Milad Shokouhi, Emine Yilmaz |
| 2012 | SIGIR | On judgments obtained from a commercial search engine. | Emine Yilmaz, Gabriella Kazai, Nick Craswell, Seyed M. M. Tahaghoghi |
| 2011 | CIKM | Simulating simple user behavior for system effectiveness evaluation. | Ben Carterette, Evangelos Kanoulas, Emine Yilmaz |
| 2011 | CIKM | Semi-supervised learning to rank with preference regularization. | Martin Szummer, Emine Yilmaz |
| 2011 | CIKM | Relevance feedback exploiting query-specific document manifolds. | Chang Wang, Emine Yilmaz, Martin Szummer |
| 2011 | SIGIR | Inferring and using location metadata to personalize web search. | Paul N. Bennett, Filip Radlinski, Ryen W. White, Emine Yilmaz |
| 2011 | WSDM | Crowdsourcing for search and data mining. | Vitor R. Carvalho, Matthew Lease, Emine Yilmaz |
| 2011 | WSDM | Detecting duplicate web documents using clickthrough data. | Filip Radlinski, Paul N. Bennett, Emine Yilmaz |
| 2010 | CIKM | Expected browsing utility for web search evaluation. | Emine Yilmaz, Milad Shokouhi, Nick Craswell, Stephen Robertson |
| 2010 | SIGIR | Low cost evaluation in information retrieval. | Ben Carterette, Evangelos Kanoulas, Emine Yilmaz |
| 2010 | SIGIR | Extending average precision to graded relevance judgments. | Stephen E. Robertson, Evangelos Kanoulas, Emine Yilmaz |
| 2009 | SIGIR | Document selection methodologies for efficient and effective learning-to-rank. | Javed A. Aslam, Evangelos Kanoulas, Virgiliu Pavlu, Stefan Savev, Emine Yilmaz |
| 2009 | SIGIR | Deep versus shallow judgments in learning to rank. | Emine Yilmaz, Stephen Robertson |
| 2009 | SIGIR | Incorporating User Behavior Information in IR Evaluation. | Emine Yilmaz, Milad Shokouhi, Nick Craswell, Stephen Robertson |
| 2008 | SIGIR | Relevance assessment: are judges exchangeable and does it matter. | Peter Bailey, Nick Craswell, Ian Soboroff, Paul Thomas, Arjen P. de Vries, Emine Yilmaz |
| 2008 | SIGIR | A new rank correlation coefficient for information retrieval. | Emine Yilmaz, Javed A. Aslam, Stephen Robertson |
| 2008 | SIGIR | A simple and efficient sampling method for estimating AP and NDCG. | Emine Yilmaz, Evangelos Kanoulas, Javed A. Aslam |
| 2007 | CIKM | Inferring document relevance from incomplete information. | Javed A. Aslam, Emine Yilmaz |
| 2006 | CIKM | Estimating average precision with incomplete and imperfect judgments. | Emine Yilmaz, Javed A. Aslam |
| 2006 | SIGIR | A statistical method for system evaluation using incomplete judgments. | Javed A. Aslam, Virgiliu Pavlu, Emine Yilmaz |
| 2006 | SIGIR | Inferring document relevance via average precision. | Javed A. Aslam, Emine Yilmaz |
| 2005 | CIKM | A geometric interpretation and analysis of R-precision. | Javed A. Aslam, Emine Yilmaz |
| 2005 | SIGIR | Measure-based metasearch. | Javed A. Aslam, Virgiliu Pavlu, Emine Yilmaz |
| 2005 | SIGIR | The maximum entropy method for analyzing retrieval measures. | Javed A. Aslam, Emine Yilmaz, Virgiliu Pavlu |
| 2005 | SIGIR | A geometric interpretation of r-precision and its correlation with average precision. | Javed A. Aslam, Emine Yilmaz, Virgiliu Pavlu |