| 2014 | ADCS | Improving test collection pools with machine learning. | Gaya K. Jayasinghe, William Webber, Mark Sanderson, J. Shane Culpepper |
| 2014 | ECIR | Reducing Reliance on Relevance Judgments for System Comparison by Using Expectation-Maximization. | Ning Gao, William Webber, Douglas W. Oard |
| 2014 | SIGIR | Extending test collection pools without manual runs. | Gaya K. Jayasinghe, William Webber, Mark Sanderson, J. Shane Culpepper |
| 2014 | SIGIR | Evaluating non-deterministic retrieval systems. | Gaya K. Jayasinghe, William Webber, Mark Sanderson, Lasitha Sandamali Dharmasena, J. Shane Culpepper |
| 2013 | CIKM | Towards minimizing the annotation cost of certified text classification. | Mossaab Bagdouri, William Webber, David D. Lewis, Douglas W. Oard |
| 2013 | SIGIR | Document features predicting assessor disagreement. | Praveen Chandar, William Webber, Ben Carterette |
| 2013 | SIGIR | The effect of threshold priming and need for cognition on relevance calibration and assessment. | Falk Scholer, Diane Kelly, Wan-Ching Wu, Hanseul S. Lee, William Webber |
| 2013 | SIGIR | Sequential testing in classifier evaluation yields biased estimates of effectiveness. | William Webber, Mossaab Bagdouri, David D. Lewis, Douglas W. Oard |
| 2013 | SIGIR | Assessor disagreement and text classifier accuracy. | William Webber, Jeremy Pickens |
| 2012 | CIKM | Alternative assessor disagreement and retrieval depth. | William Webber, Praveen Chandar, Ben Carterette |
| 2012 | SIGIR | Effect of written instructions on assessor agreement. | William Webber, Bryan Toth, Marjorie Desamito |
| 2011 | CIKM | Principles for robust evaluation infrastructure. | Justin Zobel, William Webber, Mark Sanderson, Alistair Moffat |
| 2010 | CIKM | Assessor error in stratified evaluation. | William Webber, Douglas W. Oard, Falk Scholer, Bruce Hedin |
| 2009 | CIKM | Improvements that don't add up: ad-hoc retrieval results since 1998. | Timothy G. Armstrong, Alistair Moffat, William Webber, Justin Zobel |
| 2009 | SIGIR | Has adhoc retrieval improved since 1994? | Timothy G. Armstrong, Alistair Moffat, William Webber, Justin Zobel |
| 2009 | SIGIR | EvaluatIR: an online tool for evaluating and comparing IR systems. | Timothy G. Armstrong, Alistair Moffat, William Webber, Justin Zobel |
| 2009 | SIGIR | Score adjustment for correction of pooling bias. | William Webber, Laurence Anthony F. Park |
| 2008 | CIKM | Statistical power in retrieval experimentation. | William Webber, Alistair Moffat, Justin Zobel |
| 2008 | SIGIR | Score standardization for inter-collection comparison of retrieval systems. | William Webber, Alistair Moffat, Justin Zobel |
| 2008 | SIGIR | Precision-at-ten considered redundant. | William Webber, Alistair Moffat, Justin Zobel, Tetsuya Sakai |
| 2007 | SIGIR | Strategic system comparisons via targeted relevance judgments. | Alistair Moffat, William Webber, Justin Zobel |
| 2006 | SIGIR | Load balancing for term-distributed parallel retrieval. | Alistair Moffat, William Webber, Justin Zobel |
| 2005 | WISE | Space-Limited Ranked Query Evaluation Using Adaptive Pruning. | Nicholas Lester, Alistair Moffat, William Webber, Justin Zobel |