| 2026 | ACL | BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search Agents. | Zijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie, Kai Zou, Sahel Sharifymoghaddam, Andrew Liu, Joshua Green, Kshama Patel, Ruoxi Meng, Mingyi Su, Yanxi Li, Haoran Hong, Xinyu Shi, Xuye Liu, Hosna Oyarhoseini, Nandan Thakur, Crystina Zhang, Luyu Gao, Wenhu Chen, Jimmy Lin |
| 2026 | ACL | rosaOS: Agentic Operating System for Embodied LLMs. | Yijun Ge, Kushaldeep Mujral, Karthik Nambiar, Jimmy Lin |
| 2026 | ACL | Rerank Before You Reason: Analyzing Reranking Tradeoffs through Effective Token Cost in Deep Search Agents. | Sahel Sharifymoghaddam, Jimmy Lin |
| 2026 | ECIR | Contrastive Learning Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data. | Manveer Singh Tamber, Suleman Kazi, Vivek Sourabh, Jimmy Lin |
| 2026 | ECIR | Understanding Multi-Structured Documents via LLMs'. | Shivani Upadhyay, Messiah Ataey, Syed Shariyar Murtaza, Yifan Nie, Anirudh Aggarwal, Jimmy Lin |
| 2026 | ECIR | Do We Still Need Text Features for Video Retrieval in the Era of Vision-Language Models? | Jiaqi Samantha Zhan, Crystina Zhang, Shengyao Zhuang, Xueguang Ma, Jimmy Lin |
| 2026 | ICTIR | Search Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses. | Sahel Sharifymoghaddam, Shivani Upadhyay, Nandan Thakur, Ronak Pradeep, Jimmy Lin |
| 2026 | SIGIR | MCP Servers for Pyserini and RankLLM: Enabling Agentic Retrieval-Augmented Generation. | Yijun Ge, Zibo Guo, Sahel Sharifymoghaddam, Jimmy Lin |
| 2026 | SIGIR | NanoKnow: How to Know What Your Language Model Knows. | Lingwei Gu, Nour Jedidi, Jimmy Lin |
| 2026 | SIGIR | Revisiting BM25 Feedback Models using HyDE. | Nour Jedidi, Jimmy Lin |
| 2026 | SIGIR | Lighting the Way for BRIGHT: Reproducible Baselines with Anserini, Pyserini, and RankLLM. | Sahel Sharifymoghaddam, Yijun Ge, Raghav Vasudeva, Jimmy Lin |
| 2026 | SIGIR | Automating Generation of Long-Form Queries. | Shivani Upadhyay, Daniel Campos, Nandan Thakur, Ronak Pradeep, Nick Craswell, Jimmy Lin |
| 2026 | SIGIR | LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum. | Zhichao Xu, Shengyao Zhuang, Crystina Zhang, Xueguang Ma, Yijun Tian, Maitrey Mehta, Jimmy Lin, Vivek Srikumar |
| 2025 | ACL | Operational Advice for Dense and Sparse Retrievers: HNSW, Flat, or Inverted Indexes? | Jimmy Lin |
| 2025 | ACL | DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers. | Xueguang Ma, Xi Victoria Lin, Barlas Oguz, Jimmy Lin, Wen-tau Yih, Xilun Chen |
| 2025 | ACL | VISA: Retrieval Augmented Generation with Visual Source Attribution. | Xueguang Ma, Shengyao Zhuang, Bevan Koopman, Guido Zuccon, Wenhu Chen, Jimmy Lin |
| 2025 | ACL | AfroBench: How Good are Large Language Models on African Languages? | Jessica Ojo, Odunayo Ogundepo, Akintunde Oladipo, Kelechi Ogueji, Jimmy Lin, Pontus Stenetorp, David Ifeoluwa Adelani |
| 2025 | CIKM | Study on LLMs for Promptagator-Style Dense Retriever Training. | Daniel Gwon, Nour Jedidi, Jimmy Lin |
| 2025 | ECIR | The Impact of Incidental Multilingual Text on Cross-Lingual Transfer in Monolingual Retrieval. | Andrew Liu, Edward Xu, Crystina Zhang, Jimmy Lin |
| 2025 | ECIR | Ragnark: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track. | Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, Jimmy Lin |
| 2025 | ECIR | LiT and Lean: Distilling Listwise Rerankers Into Encoder-Decoder Models. | Manveer Singh Tamber, Ronak Pradeep, Jimmy Lin |
| 2025 | ECIR | Patience in Proximity: A Simple Early Termination Strategy for HNSW Graph Traversal in Approximate k-Nearest Neighbor Search. | Tommaso Teofili, Jimmy Lin |
| 2025 | ECIR | Rank-Without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models. | Crystina Zhang, Sebastian Hofsttter, Patrick Lewis, Raphael Tang, Jimmy Lin |
| 2025 | EMNLP | QuackIR: Retrieval in DuckDB and Other Relational Database Management Systems. | Yijun Ge, Zijian Chen, Jimmy Lin |
| 2025 | EMNLP | Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards. | Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu, Ge Luo, Suleman Kazi, Minseok Bae, Miaoran Li, Ofer Mendelevitch, Renyi Qu, Jimmy Lin |
| 2025 | EMNLP | Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs. | Nandan Thakur, Crystina Zhang, Xueguang Ma, Jimmy Lin |
| 2025 | ICLR | Mm-Embed: Universal Multimodal Retrieval with Multimodal LLMS. | Sheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin, Bryan Catanzaro, Wei Ping |
| 2025 | IJCNLP | Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges. | Manveer Singh Tamber, Jimmy Lin |
| 2025 | ICTIR | A Large-Scale Study of Relevance Assessments with Large Language Models Using UMBRELA. | Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Daniel Campos, Nick Craswell, Ian Soboroff, Jimmy Lin |
| 2025 | KDD | CURE: A dataset for Clinical Understanding & Retrieval Evaluation. | Nadia Sheikh, Daniel Buades Marcos, Anne-Laure Jousse, Akintunde Oladipo, Olivier Rousseau, Jimmy Lin |
| 2025 | NAACL | Zero-Shot ATC Coding with Large Language Models for Clinical Assessments. | Zijian Chen, John-Michael Gamble, Jimmy Lin |
| 2025 | NAACL | UniRAG: Universal Retrieval Augmentation for Large Vision Language Models. | Sahel Sharifymoghaddam, Shivani Upadhyay, Wenhu Chen, Jimmy Lin |
| 2025 | NAACL | Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models. | Manveer Singh Tamber, Jasper Xian, Jimmy Lin |
| 2025 | NAACL | MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems. | Nandan Thakur, Suleman Kazi, Ge Luo, Jimmy Lin, Amin Ahmad |
| 2025 | NAACL | Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts? | Crystina Zhang, Jing Lu, Vinh Q. Tran, Tal Schuster, Donald Metzler, Jimmy Lin |
| 2025 | SIGIR | Accelerating Listwise Reranking: Reproducing and Enhancing FIRST. | Zijian Chen, Ronak Pradeep, Jimmy Lin |
| 2025 | SIGIR | Gosling Grows Up: Retrieval with Learned Dense and Sparse Representations Using Anserini. | Jimmy Lin, Arthur Haonan Chen, Carlos Lassance, Xueguang Ma, Ronak Pradeep, Tommaso Teofili, Jasper Xian, Jheng-Hong Yang, Brayden Zhong, Vincent Zhong |
| 2025 | SIGIR | Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality. | Xueguang Ma, Luyu Gao, Shengyao Zhuang, Jiaqi Samantha Zhan, Jamie Callan, Jimmy Lin |
| 2025 | SIGIR | The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models. | Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, Ian Soboroff, Hoa Trang Dang, Jimmy Lin |
| 2025 | SIGIR | RankLLM: A Python Package for Reranking with LLMs. | Sahel Sharifymoghaddam, Ronak Pradeep, Andre Slavescu, Ryan Nguyen, Andrew Xu, Zijian Chen, Yilin Zhang, Yidi Chen, Jasper Xian, Jimmy Lin |
| 2025 | SIGIR | Assessing Support for the TREC 2024 RAG Track: A Large-Scale Comparative Study of LLM and Human Evaluations. | Nandan Thakur, Ronak Pradeep, Shivani Upadhyay, Daniel Campos, Nick Craswell, Ian Soboroff, Hoa Trang Dang, Jimmy Lin |
| 2024 | AAAI | Jointly Modeling Spatio-Temporal Features of Tactile Signals for Action Classification. | Jimmy Lin, Junkai Li, Jiasi Gao, Weizhi Ma, Yang Liu |
| 2024 | ACL | Zero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages. | Mofetoluwa Adeyemi, Akintunde Oladipo, Ronak Pradeep, Jimmy Lin |
| 2024 | ACL | EWEK-QA : Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering Systems. | Mohammad Dehghan, Mohammad Ali Alomrani, Sunyam Bagga, David Alfonso-Hermelo, Khalil Bibi, Abbas Ghaddar, Yingxue Zhang, Xiaoguang Li, Jianye Hao, Qun Liu, Jimmy Lin, Boxing Chen, Prasanna Parthasarathi, Mahdi Biparva, Mehdi Rezagholizadeh |
| 2024 | ECIR | Towards Automated End-to-End Health Misinformation Free Search with a Large Language Model. | Ronak Pradeep, Jimmy Lin |
| 2024 | EMNLP | Unifying Multimodal Retrieval via Document Screenshot Embedding. | Xueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen, Jimmy Lin |
| 2024 | EMNLP | ConvKGYarn: Spinning Configurable and Scalable Conversational Knowledge Graph QA Datasets with Large Language Models. | Ronak Pradeep, Daniel Lee, Ali Mousavi, Jeffrey Pound, Yisi Sang, Jimmy Lin, Ihab F. Ilyas, Saloni Potdar, Mostafa Arefiyan, Yunyao Li |
| 2024 | EMNLP | Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation. | Raphael Tang, Xinyu Zhang, Lixinyu Xu, Yao Lu, Wenyan Li, Pontus Stenetorp, Jimmy Lin, Ferhan Ture |
| 2024 | EMNLP | "Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation. | Nandan Thakur, Luiz Bonifacio, Xinyu Zhang, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Boxing Chen, Mehdi Rezagholizadeh, Jimmy Lin |
| 2024 | EMNLP | PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval. | Shengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin, Guido Zuccon |
| 2024 | NAACL | Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language Models. | Raphael Tang, Xinyu Zhang, Xueguang Ma, Jimmy Lin, Ferhan Ture |
| 2024 | NAACL | Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval. | Nandan Thakur, Jianmo Ni, Gustavo Hernndez brego, John Wieting, Jimmy Lin, Daniel Cer |
| 2024 | NAACL | CELI: Simple yet Effective Approach to Enhance Out-of-Domain Generalization of Cross-Encoders. | Xinyu Zhang, Minghan Li, Jimmy Lin |
| 2024 | SIGIR | Can Query Expansion Improve Generalization of Strong Cross-Encoder Rankers? | Minghan Li, Honglei Zhuang, Kai Hui, Zhen Qin, Jimmy Lin, Rolf Jagerman, Xuanhui Wang, Michael Bendersky |
| 2024 | SIGIR | CIRAL: A Test Collection for CLIR Evaluations in African Languages. | Mofetoluwa Adeyemi, Akintunde Oladipo, Xinyu Zhang, David Alfonso-Hermelo, Mehdi Rezagholizadeh, Boxing Chen, Abdul-Hakeem Omotayo, Idris Abdulmumin, Naome A. Etori, Toyib Babatunde Musa, Samuel Fanijo, Oluwabusayo Olufunke Awoyomi, Saheed Abdullahi Salahudeen, Labaran Adamu Mohammed, Daud Olamide Abolade, Falalu Ibrahim Lawan, Maryam Sabo Abubakar, Ruqayya Nasir Iro, Amina Abubakar Imam, Shafie Abdi Mohamed, Hanad Mohamud Mohamed, Tunde Oluwaseyi Ajayi, Jimmy Lin |
| 2024 | SIGIR | Resources for Brewing BEIR: Reproducible Reference Models and Statistical Analyses. | Ehsan Kamalloo, Nandan Thakur, Carlos Lassance, Xueguang Ma, Jheng-Hong Yang, Jimmy Lin |
| 2024 | SIGIR | Towards Robust QA Evaluation via Open LLMs. | Ehsan Kamalloo, Shivani Upadhyay, Jimmy Lin |
| 2024 | SIGIR | Fine-Tuning LLaMA for Multi-Stage Text Retrieval. | Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, Jimmy Lin |
| 2024 | SIGIR | On Backbones and Training Regimes for Dense Retrieval in African Languages. | Akintunde Oladipo, Mofetoluwa Adeyemi, Jimmy Lin |
| 2024 | SIGIR | Systematic Evaluation of Neural Retrieval Models on the Touch 2020 Argument Retrieval Subset of BEIR. | Nandan Thakur, Luiz Bonifacio, Maik Frbe, Alexander Bondarenko, Ehsan Kamalloo, Martin Potthast, Matthias Hagen, Jimmy Lin |
| 2023 | ACL | Precise Zero-Shot Dense Retrieval without Relevance Labels. | Luyu Gao, Xueguang Ma, Jimmy Lin, Jamie Callan |
| 2023 | ACL | "Low-Resource" Text Classification: A Parameter-Free Classification Method with Compressors. | Zhiying Jiang, Matthew Y. R. Yang, Mikhail Tsirlin, Raphael Tang, Yiqin Dai, Jimmy Lin |
| 2023 | ACL | Evaluating Embedding APIs for Information Retrieval. | Ehsan Kamalloo, Xinyu Zhang, Odunayo Ogundepo, Nandan Thakur, David Alfonso-Hermelo, Mehdi Rezagholizadeh, Jimmy Lin |
| 2023 | ACL | CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval. | Minghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, Xilun Chen |
| 2023 | ACL | GAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Exploration. | Aleksandra Piktus, Odunayo Ogundepo, Christopher Akiki, Akintunde Oladipo, Xinyu Zhang, Hailey Schoelkopf, Stella Biderman, Martin Potthast, Jimmy Lin |
| 2023 | ACL | What the DAAM: Interpreting Stable Diffusion Using Cross Attention. | Raphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang, Gefei Yang, Karun Kumar, Pontus Stenetorp, Jimmy Lin, Ferhan Ture |
| 2023 | ACL | Operator Selection and Ordering in a Pipeline Approach to Efficiency Optimizations for Transformers. | Ji Xin, Raphael Tang, Zhiying Jiang, Yaoliang Yu, Jimmy Lin |
| 2023 | CIKM | Anserini Gets Dense Retrieval: Integration of Lucene's HNSW Indexes. | Xueguang Ma, Tommaso Teofili, Jimmy Lin |
| 2023 | ECIR | PyGaggle: A Gaggle of Resources for Open-Domain Question Answering. | Ronak Pradeep, Haonan Chen, Lingwei Gu, Manveer Singh Tamber, Jimmy Lin |
| 2023 | ECIR | Pre-processing Matters! Improved Wikipedia Corpora for Open-Domain Question Answering. | Manveer Singh Tamber, Ronak Pradeep, Jimmy Lin |
| 2023 | EMNLP | Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face. | Christopher Akiki, Odunayo Ogundepo, Aleksandra Piktus, Xinyu Zhang, Akintunde Oladipo, Jimmy Lin, Martin Potthast |
| 2023 | EMNLP | mAggretriever: A Simple yet Effective Approach to Zero-Shot Multilingual Dense Retrieval. | Sheng-Chieh Lin, Amin Ahmad, Jimmy Lin |
| 2023 | EMNLP | How to Train Your Dragon: Diverse Augmentation Towards Generalizable Dense Retrieval. | Sheng-Chieh Lin, Akari Asai, Minghan Li, Barlas Oguz, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, Xilun Chen |
| 2023 | EMNLP | Better Quality Pre-training Data and T5 Models for African Languages. | Akintunde Oladipo, Mofetoluwa Adeyemi, Orevaoghene Ahia, Abraham Toluwase Owodunni, Odunayo Ogundepo, David Ifeoluwa Adelani, Jimmy Lin |
| 2023 | EMNLP | How Does Generative Retrieval Scale to Millions of Passages? | Ronak Pradeep, Kai Hui, Jai Gupta, dm D. Lelkes, Honglei Zhuang, Jimmy Lin, Donald Metzler, Vinh Q. Tran |
| 2023 | SIGIR | Tevatron: An Efficient and Flexible Toolkit for Neural Retrieval. | Luyu Gao, Xueguang Ma, Jimmy Lin, Jamie Callan |
| 2023 | SIGIR | MMEAD: MS MARCO Entity Annotations and Disambiguations. | Chris Kamphuis, Aileen Lin, Siwen Yang, Jimmy Lin, Arjen P. de Vries, Faegheh Hasibi |
| 2023 | SIGIR | SLIM: Sparsified Late Interaction for Multi-Vector Retrieval with Inverted Indexes. | Minghan Li, Sheng-Chieh Lin, Xueguang Ma, Jimmy Lin |
| 2023 | SIGIR | SPRINT: A Unified Toolkit for Evaluating and Demystifying Zero-shot Neural Sparse Retrieval. | Nandan Thakur, Kexin Wang, Iryna Gurevych, Jimmy Lin |
| 2023 | SIGIR | AToMiC: An Image/Text Retrieval Test Collection to Support Multimedia Content Creation. | Jheng-Hong Yang, Carlos Lassance, Rafael Sampaio de Rezende, Krishna Srinivasan, Miriam Redi, Stphane Clinchant, Jimmy Lin |
| 2022 | ADCS | Pseudo-Relevance Feedback with Dense Retrievers in Pyserini. | Hang Li, Shengyao Zhuang, Xueguang Ma, Jimmy Lin, Guido Zuccon |
| 2022 | ECIR | Improving Query Representations for Dense Retrieval with Pseudo Relevance Feedback: A Reproducibility Study. | Hang Li, Shengyao Zhuang, Ahmed Mourad, Xueguang Ma, Jimmy Lin, Guido Zuccon |
| 2022 | ECIR | Another Look at DPR: Reproduction of Training and Replication of Retrieval. | Xueguang Ma, Kai Sun, Ronak Pradeep, Minghan Li, Jimmy Lin |
| 2022 | ECIR | Squeezing Water from a Stone: A Bag of Tricks for Further Improving Cross-Encoder Effectiveness for Reranking. | Ronak Pradeep, Yuqi Liu, Xinyu Zhang, Yilin Li, Andrew Yates, Jimmy Lin |
| 2022 | EMNLP | Cross-lingual Text-to-SQL Semantic Parsing with Representation Mixup. | Peng Shi, Linfeng Song, Lifeng Jin, Haitao Mi, He Bai, Jimmy Lin, Dong Yu |
| 2022 | EMNLP | XRICL: Cross-lingual Retrieval-Augmented In-Context Learning for Cross-lingual Text-to-SQL Semantic Parsing. | Peng Shi, Rui Zhang, He Bai, Jimmy Lin |
| 2022 | EMNLP | Certified Error Control of Candidate Set Pruning for Two-Stage Relevance Ranking. | Minghan Li, Xinyu Zhang, Ji Xin, Hongyang Zhang, Jimmy Lin |
| 2022 | EMNLP | AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages. | Odunayo Ogundepo, Xinyu Zhang, Shuo Sun, Kevin Duh, Jimmy Lin |
| 2022 | EMNLP | SpeechNet: Weakly Supervised, End-to-End Speech Recognition at Industrial Scale. | Raphael Tang, Karun Kumar, Gefei Yang, Akshat Pandey, Yajie Mao, Vladislav Belyaev, Madhuri Emmadi, G. Craig Murray, Ferhan Ture, Jimmy Lin |
| 2022 | EMNLP | Improving Precancerous Case Characterization via Transformer-based Ensemble Learning. | Yizhen Zhong, Jiajie Xiao, Thomas Vetterli, Mahan Matin, Ellen Loo, Jimmy Lin, Richard Bourgon, Ofer Shapira |
| 2022 | EMNLP | Evaluating Token-Level and Passage-Level Dense Retrieval Models for Math Information Retrieval. | Wei Zhong, Jheng-Hong Yang, Yuqing Xie, Jimmy Lin |
| 2022 | ICASSP | Temporal Early Exiting for Streaming Speech Commands Recognition. | Raphael Tang, Karun Kumar, Ji Xin, Piyush Vyas, Wenyan Li, Gefei Yang, Yajie Mao, G. Craig Murray, Jimmy Lin |
| 2022 | SIGIR | Fostering Coopetition While Plugging Leaks: The Design and Implementation of the MS MARCO Leaderboards. | Jimmy Lin, Daniel Campos, Nick Craswell, Bhaskar Mitra, Emine Yilmaz |
| 2022 | SIGIR | Another Look at Information Retrieval as Statistical Translation. | Yuqi Liu, Chengcheng Hu, Jimmy Lin |
| 2022 | SIGIR | To Interpolate or not to Interpolate: PRF, Dense and Sparse Retrievers. | Hang Li, Shuai Wang, Shengyao Zhuang, Ahmed Mourad, Xueguang Ma, Jimmy Lin, Guido Zuccon |
| 2022 | SIGIR | Document Expansion Baselines and Learned Sparse Lexical Representations for MS MARCO V1 and V2. | Xueguang Ma, Ronak Pradeep, Rodrigo Nogueira, Jimmy Lin |
| 2022 | SIGIR | Neural Query Synthesis and Domain-Specific Ranking Templates for Multi-Stage Clinical Trial Matching. | Ronak Pradeep, Yilin Li, Yuetong Wang, Jimmy Lin |
| 2022 | SIGIR | Flipping the Script: Inverse Information Seeking Dialogues for Market Research. | Josh Seltzer, Kathy Cheng, Shi Zong, Jimmy Lin |
| 2022 | SIGIR | A Common Framework for Exploring Document-at-a-Time and Score-at-a-Time Retrieval Methods. | Andrew Trotman, Joel Mackenzie, Pradeesh Parameswaran, Jimmy Lin |
| 2022 | SIGIR | Too Many Relevants: Whither Cranfield Test Collections? | Ellen M. Voorhees, Nick Craswell, Jimmy Lin |
| 2021 | AAAI | Segatron: Segment-Aware Transformer for Language Modeling and Understanding. | He Bai, Peng Shi, Jimmy Lin, Yuqing Xie, Luchen Tan, Kun Xiong, Wen Gao, Ming Li |
| 2021 | ACL | Semantics of the Unwritten: The Effect of End of Paragraph and Sequence Tokens on Text Generation with GPT2. | He Bai, Peng Shi, Jimmy Lin, Luchen Tan, Kun Xiong, Wen Gao, Jie Liu, Ming Li |
| 2021 | ACL | Exploring Listwise Evidence Reasoning with T5 for Fact Verification. | Kelvin Jiang, Ronak Pradeep, Jimmy Lin |
| 2021 | ACL | The Art of Abstention: Selective Prediction and Error Regularization for Natural Language Processing. | Ji Xin, Raphael Tang, Yaoliang Yu, Jimmy Lin |
| 2021 | DocEng | Rescuing historical climate observations to support hydrological research: a case study of solar radiation data. | Ogundepo Odunayo, Naveela N. Sookoo, Gautam Bathla, Anthony Cavallin, Bhaleka D. Persaud, Kathy Szigeti, Philippe Van Cappellen, Jimmy Lin |
| 2021 | EACL | BERxiT: Early Exiting for BERT with Better Fine-Tuning and Extension to Regression. | Ji Xin, Raphael Tang, Yaoliang Yu, Jimmy Lin |
| 2021 | EACL | Don't Change Me! User-Controllable Selective Paraphrase Generation. | Mohan Zhang, Luchen Tan, Zihang Fu, Kun Xiong, Jimmy Lin, Ming Li, Zhengkai Tu |
| 2021 | ECIR | Comparing Score Aggregation Approaches for Document Retrieval with Pretrained Transformers. | Xinyu Zhang, Andrew Yates, Jimmy Lin |
| 2021 | EMNLP | Unsupervised Chunking as Syntactic Structure Induction with a Knowledge-Transfer Approach. | Anup Anand Deshmukh, Qianqiu Zhang, Ming Li, Jimmy Lin, Lili Mou |
| 2021 | EMNLP | Multi-Task Dense Retrieval via Model Uncertainty Fusion for Open-Domain Question Answering. | Minghan Li, Ming Li, Kun Xiong, Jimmy Lin |
| 2021 | EMNLP | Contextualized Query Embeddings for Conversational Search. | Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin |
| 2021 | EMNLP | Simple and Effective Unsupervised Redundancy Elimination to Compress Dense Vectors for Passage Retrieval. | Xueguang Ma, Minghan Li, Kai Sun, Ji Xin, Jimmy Lin |
| 2021 | EMNLP | Voice Query Auto Completion. | Raphael Tang, Karun Kumar, Kendra Chalkley, Ji Xin, Liming Zhang, Wenyan Li, Gefei Yang, Yajie Mao, Junho Shin, Geoffrey Craig Murray, Jimmy Lin |
| 2021 | EMNLP | Learning to Rank in the Age of Muppets: Effectiveness-Efficiency Tradeoffs in Multi-Stage Ranking. | Yue Zhang, Chengcheng Hu, Yuqi Liu, Hui Fang, Jimmy Lin |
| 2021 | ICTIR | The Simplest Thing That Can Possibly Work: (Pseudo-)Relevance Feedback via Text Classification. | Xiao Han, Yuqi Liu, Jimmy Lin |
| 2021 | SIGIR | MS MARCO: Benchmarking Ranking Models in the Large-Data Regime. | Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Jimmy Lin |
| 2021 | SIGIR | Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware Sampling. | Sebastian Hofsttter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, Allan Hanbury |
| 2021 | SIGIR | Significant Improvements over the State of the Art? A Case Study of the MS MARCO Document Ranking Leaderboard. | Jimmy Lin, Daniel Campos, Nick Craswell, Bhaskar Mitra, Emine Yilmaz |
| 2021 | SIGIR | Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations. | Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, Rodrigo Nogueira |
| 2021 | SIGIR | Vera: Prediction Techniques for Reducing Harmful Misinformation in Consumer Health Search. | Ronak Pradeep, Xueguang Ma, Rodrigo Nogueira, Jimmy Lin |
| 2021 | SIGIR | Pretrained Transformers for Text Ranking: BERT and Beyond. | Andrew Yates, Rodrigo Nogueira, Jimmy Lin |
| 2021 | SIGIR | Chatty Goose: A Python Framework for Conversational Search. | Edwin Zhang, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, Rodrigo Nogueira, Jimmy Lin |
| 2020 | ACL | Two Birds, One Stone: A Simple, Unified Model for Text Generation from Structured and Unstructured Data. | Hamidreza Shahidi, Ming Li, Jimmy Lin |
| 2020 | ACL | Showing Your Work Doesn't Always Work. | Raphael Tang, Jaejun Lee, Ji Xin, Xinyu Liu, Yaoliang Yu, Jimmy Lin |
| 2020 | ACL | DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference. | Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, Jimmy Lin |
| 2020 | CHIIR | We Could, but Should We?: Ethical Considerations for Providing Access to GeoCities and Other Historical Digital Collections. | Jimmy Lin, Ian Milligan, Douglas W. Oard, Nick Ruest, Katie Shilton |
| 2020 | CHIIR | Update Delivery Mechanisms for Prospective Information Needs: A Reproducibility Study. | Royal Sequiera, Luchen Tan, Yinan Zhang, Jimmy Lin |
| 2020 | CIKM | Flexible IR Pipelines with Capreolus. | Andrew Yates, Kevin Martin Jose, Xinyu Zhang, Jimmy Lin |
| 2020 | COLING | Designing Templates for Eliciting Commonsense Knowledge from Pretrained Sequence-to-Sequence Models. | Jheng-Hong Yang, Sheng-Chieh Lin, Rodrigo Nogueira, Ming-Feng Tsai, Chuan-Ju Wang, Jimmy Lin |
| 2020 | ECIR | From MAXSCORE to Block-Max Wand: The Story of How Lucene Significantly Improved Query Evaluation Performance. | Adrien Grand, Robert Muir, Jim Ferenczi, Jimmy Lin |
| 2020 | ECIR | Which BM25 Do You Mean? A Large-Scale Reproducibility Study of Scoring Variants. | Chris Kamphuis, Arjen P. de Vries, Leonid Boytsov, Jimmy Lin |
| 2020 | ECIR | Reproducibility is a Process, Not an Achievement: The Replicability of IR Reproducibility Experiments. | Jimmy Lin, Qian Zhang |
| 2020 | EMNLP | Cydex: Neural Search Infrastructure for the Scholarly Literature. | Shane Ding, Edwin Zhang, Jimmy Lin |
| 2020 | EMNLP | Inserting Information Bottleneck for Attribution in Transformers. | Zhiying Jiang, Raphael Tang, Ji Xin, Jimmy Lin |
| 2020 | EMNLP | Document Ranking with a Pretrained Sequence-to-Sequence Model. | Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, Jimmy Lin |
| 2020 | EMNLP | Cross-Lingual Training of Neural Models for Document Ranking. | Peng Shi, He Bai, Jimmy Lin |
| 2020 | EMNLP | Early Exiting BERT for Efficient Document Ranking. | Ji Xin, Rodrigo Nogueira, Yaoliang Yu, Jimmy Lin |
| 2020 | EMNLP | Covidex: Neural Ranking Models and Keyword Search Infrastructure for the COVID-19 Open Research Dataset. | Edwin Zhang, Nikhil Gupta, Raphael Tang, Xiao Han, Ronak Pradeep, Kuang Lu, Yue Zhang, Rodrigo Nogueira, Kyunghyun Cho, Hui Fang, Jimmy Lin |
| 2020 | EMNLP | A Little Bit Is Worse Than None: Ranking with Limited Training Data. | Xinyu Zhang, Andrew Yates, Jimmy Lin |
| 2020 | ICML | Generalized and Scalable Optimal Sparse Decision Trees. | Jimmy Lin, Chudi Zhong, Diane Hu, Cynthia Rudin, Margo I. Seltzer |
| 2020 | ICTIR | Approximate Nearest Neighbor Search and Lightweight Dense Vector Reranking in Multi-Stage Retrieval Architectures. | Zhengkai Tu, Wei Yang, Zihang Fu, Yuqing Xie, Luchen Tan, Kun Xiong, Ming Li, Jimmy Lin |
| 2020 | KDD | SimClusters: Community-Based Representations for Heterogeneous Recommendations at Twitter. | Venu Satuluri, Yao Wu, Xun Zheng, Yilei Qian, Brian Wichers, Qieyun Dai, Gui Ming Tang, Jerry Jiang, Jimmy Lin |
| 2020 | WWW | Distant Supervision for Multi-Stage Fine-Tuning in Retrieval-Based Question Answering. | Yuqing Xie, Wei Yang, Luchen Tan, Kun Xiong, Nicholas Jing Yuan, Baoxing Huai, Ming Li, Jimmy Lin |
| 2020 | SIGIR | Supporting Interoperability Between Open-Source Search Engines with the Common Index File Format. | Jimmy Lin, Joel M. Mackenzie, Chris Kamphuis, Craig Macdonald, Antonio Mallia, Michal Siedlaczek, Andrew Trotman, Arjen P. de Vries |
| 2020 | SIGIR | A Lightweight Environment for Learning Experimental IR Research Practices. | Zeynep Akkalyoncu Yilmaz, Charles L. A. Clarke, Jimmy Lin |
| 2019 | AAAI | Multi-Perspective Relevance Matching with Hierarchical ConvNets for Social Media Search. | Jinfeng Rao, Wei Yang, Yuhao Zhang, Ferhan Tre, Jimmy Lin |
| 2019 | AMIA | Identification and Ranking of Biomedical Informatics Researcher Citation Statistics through a Google Scholar Scraper. | Allison B. McCoy, Dean F. Sittig, Jimmy Lin, Adam Wright |
| 2019 | ECIR | Reproducing and Generalizing Semantic Term Matching in Axiomatic Information Retrieval. | Peilin Yang, Jimmy Lin |
| 2019 | ECIR | Simple Techniques for Cross-Collection Relevance Feedback. | Ruifan Yu, Yuhao Xie, Jimmy Lin |
| 2019 | EMNLP | Honkling: In-Browser Personalization for Ubiquitous Keyword Spotting. | Jaejun Lee, Raphael Tang, Jimmy Lin |
| 2019 | EMNLP | Incorporating Contextual and Syntactic Structures Improves Semantic Similarity Modeling. | Linqing Liu, Wei Yang, Jinfeng Rao, Raphael Tang, Jimmy Lin |
| 2019 | EMNLP | Bridging the Gap between Relevance Matching and Semantic Matching for Short Text Similarity Modeling. | Jinfeng Rao, Linqing Liu, Yi Tay, Hsiu-Wei Yang, Peng Shi, Jimmy Lin |
| 2019 | EMNLP | What Part of the Neural Network Does This? Understanding LSTMs by Measuring and Dissecting Neurons. | Ji Xin, Jimmy Lin, Yaoliang Yu |
| 2019 | EMNLP | Aligning Cross-Lingual Entities with Multi-Aspect Information. | Hsiu-Wei Yang, Yanyan Zou, Peng Shi, Wei Lu, Jimmy Lin, Xu Sun |
| 2019 | EMNLP | Applying BERT to Document Retrieval with Birch. | Zeynep Akkalyoncu Yilmaz, Shengjin Wang, Wei Yang, Haotian Zhang, Jimmy Lin |
| 2019 | EMNLP | Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval. | Zeynep Akkalyoncu Yilmaz, Wei Yang, Haotian Zhang, Jimmy Lin |
| 2019 | IUI | Universal voice-enabled user interfaces using JavaScript. | Jaejun Lee, Raphael Tang, Jimmy Lin |
| 2019 | NAACL | Rethinking Complex Neural Network Architectures for Document Classification. | Ashutosh Adhikari, Achyudh Ram, Raphael Tang, Jimmy Lin |
| 2019 | NAACL | Simple Attention-Based Representation Learning for Ranking Short Social Media Posts. | Peng Shi, Jinfeng Rao, Jimmy Lin |
| 2019 | NAACL | Detecting Customer Complaint Escalation with Recurrent Neural Networks and Manually-Engineered Features. | Wei Yang, Luchen Tan, Chunwei Lu, Anqi Cui, Han Li, Xi Chen, Kun Xiong, Muzi Wang, Ming Li, Jian Pei, Jimmy Lin |
| 2019 | NAACL | End-to-End Open-Domain Question Answering with BERTserini. | Wei Yang, Yuqing Xie, Aileen Lin, Xingyu Li, Luchen Tan, Kun Xiong, Ming Li, Jimmy Lin |
| 2019 | SIGIR | The SIGIR 2019 Open-Source IR Replicability Challenge (OSIRRC 2019). | Ryan Clancy, Nicola Ferro, Claudia Hauff, Jimmy Lin, Tetsuya Sakai, Ze Zhong Wu |
| 2019 | SIGIR | Overview of the 2019 Open-Source IR Replicability Challenge (OSIRRC 2019). | Ryan Clancy, Nicola Ferro, Claudia Hauff, Jimmy Lin, Tetsuya Sakai, Ze Zhong Wu |
| 2019 | SIGIR | Solr Integration in the Anserini Information Retrieval Toolkit. | Ryan Clancy, Toke Eskildsen, Nick Ruest, Jimmy Lin |
| 2019 | SIGIR | Information Retrieval Meets Scalable Text Analytics: Solr Integration with Spark. | Ryan Clancy, Jaejun Lee, Zeynep Akkalyoncu Yilmaz, Jimmy Lin |
| 2019 | SIGIR | University of Waterloo Docker Images for OSIRRC at SIGIR 2019. | Ryan Clancy, Zeynep Akkalyoncu Yilmaz, Ze Zhong Wu, Jimmy Lin |
| 2019 | SIGIR | The Impact of Score Ties on Repeatability in Document Ranking. | Jimmy Lin, Peilin Yang |
| 2019 | SIGIR | Yelling at Your TV: An Analysis of Speech Recognition Errors and Subsequent User Behavior on Entertainment Systems. | Raphael Tang, Ferhan Tre, Jimmy Lin |
| 2019 | SIGIR | Challenges and Opportunities in Understanding Spoken Queries Directed at Modern Entertainment Platforms. | Ferhan Tre, Jinfeng Rao, Raphael Tang, Jimmy Lin |
| 2019 | SIGIR | Critically Examining the "Neural Hype": Weak Baselines and the Additivity of Effectiveness Gains from Neural Ranking Models. | Wei Yang, Kuang Lu, Peilin Yang, Jimmy Lin |
| 2018 | COLING | Farewell Freebase: Migrating the SimpleQuestions Dataset to DBpedia. | Michael Azmy, Peng Shi, Jimmy Lin, Ihab F. Ilyas |
| 2018 | ICASSP | Deep Residual Learning for Small-Footprint Keyword Spotting. | Raphael Tang, Jimmy Lin |
| 2018 | ICASSP | An Experimental Analysis of the Power Consumption of Convolutional Neural Networks for Keyword Spotting. | Raphael Tang, Weijie Wang, Zhucheng Tu, Jimmy Lin |
| 2018 | KDD | Multi-Task Learning with Neural Networks for Voice Query Understanding on an Entertainment Platform. | Jinfeng Rao, Ferhan Tre, Jimmy Lin |
| 2018 | NAACL | CNNs for NLP in the Browser: Client-Side Deployment and Visualization Opportunities. | Yiyun Liang, Zhucheng Tu, Laetitia Huang, Jimmy Lin |
| 2018 | NAACL | Strong Baselines for Simple Question Answering over Knowledge Graphs with and without Neural Networks. | Salman Mohammed, Peng Shi, Jimmy Lin |
| 2018 | NAACL | Pay-Per-Request Deployment of Neural Network Models Using Serverless Architectures. | Zhucheng Tu, Mengping Li, Jimmy Lin |
| 2018 | SIGIR | The Evolution of Content Analysis for Personalized Recommendations at Twitter. | Ajeet Grewal, Jimmy Lin |
| 2018 | SIGIR | Update Delivery Mechanisms for Prospective Information Needs: An Analysis of Attention in Mobile Users. | Jimmy Lin, Salman Mohammed, Royal Sequiera, Luchen Tan |
| 2018 | SIGIR | What Do Viewers Say to Their TVs?: An Analysis of Voice Queries to Entertainment Systems. | Jinfeng Rao, Ferhan Tre, Jimmy Lin |
| 2017 | CIKM | A Comparison of Nuggets and Clusters for Evaluating Timeline Summaries. | Gaurav Baruah, Richard McCreadie, Jimmy Lin |
| 2017 | CIKM | Talking to Your TV: Context-Aware Voice Search with Hierarchical Recurrent Neural Networks. | Jinfeng Rao, Ferhan Tre, Hua He, Oliver Jojic, Jimmy Lin |
| 2017 | EMNLP | An Insight Extraction System on BioMedical Literature with Deep Neural Networks. | Hua He, Kris Ganjam, Navendu Jain, Jessica Lundin, Ryen White, Jimmy Lin |
| 2017 | ICTIR | The Pareto Frontier of Utility Models as a Framework for Evaluating Push Notification Systems. | Gaurav Baruah, Jimmy Lin |
| 2017 | ICTIR | An Exploration of Serverless Architectures for Information Retrieval. | Matt Crane, Jimmy Lin |
| 2017 | ICTIR | Quantization in Append-Only Collections. | Salman Mohammed, Matt Crane, Jimmy Lin |
| 2017 | ICTIR | Mining the Temporal Statistics of Query Terms for Searching Social Media Posts. | Jinfeng Rao, Ferhan Tre, Xing Niu, Jimmy Lin |
| 2017 | WWW | Ten Blue Links on Mars. | Charles L. A. Clarke, Gordon V. Cormack, Jimmy Lin, Adam Roegiest |
| 2017 | SIGIR | The Lucene for Information Access and Retrieval Research (LIARR) Workshop at SIGIR 2017. | Leif Azzopardi, Matt Crane, Hui Fang, Grant Ingersoll, Jimmy Lin, Yashar Moshfeghi, Harrisen Scells, Peilin Yang, Guido Zuccon |
| 2017 | SIGIR | Event Detection on Curated Tweet Streams. | Nimesh Ghelani, Salman Mohammed, Shine Wang, Jimmy Lin |
| 2017 | SIGIR | Experiments with Convolutional Neural Network Models for Answer Selection. | Jinfeng Rao, Hua He, Jimmy Lin |
| 2017 | SIGIR | Online In-Situ Interleaved Evaluation of Real-Time Push Notification Systems. | Adam Roegiest, Luchen Tan, Jimmy Lin |
| 2017 | SIGIR | Finally, a Downloadable Test Collection of Tweets. | Royal Sequiera, Jimmy Lin |
| 2017 | SIGIR | On the Reusability of "Living Labs" Test Collections: : A Case Study of Real-Time Summarization. | Luchen Tan, Gaurav Baruah, Jimmy Lin |
| 2017 | SIGIR | Anserini: Enabling the Use of Lucene for Information Retrieval Research. | Peilin Yang, Hui Fang, Jimmy Lin |
| 2017 | SIGIR | Automatically Extracting High-Quality Negative Examples for Answer Selection in Question Answering. | Haotian Zhang, Jinfeng Rao, Jimmy Lin, Mark D. Smucker |
| 2016 | ADCS | Dynamic Cutoff Prediction in Multi-Stage Retrieval Systems. | J. Shane Culpepper, Charles L. A. Clarke, Jimmy Lin |
| 2016 | ADCS | In Vacuo and In Situ Evaluation of SIMD Codecs. | Andrew Trotman, Jimmy Lin |
| 2016 | CIKM | Optimizing Nugget Annotations with Active Learning. | Gaurav Baruah, Haotian Zhang, Rakesh Guttikonda, Jimmy Lin, Mark D. Smucker, Olga Vechtomova |
| 2016 | CIKM | Noise-Contrastive Estimation for Answer Selection with Deep Neural Networks. | Jinfeng Rao, Hua He, Jimmy Lin |
| 2016 | ECIR | Toward Reproducible Baselines: The Open-Source IR Reproducibility Challenge. | Jimmy Lin, Matt Crane, Andrew Trotman, Jamie Callan, Ishan Chattopadhyaya, John Foley, Grant Ingersoll, Craig MacDonald, Sebastiano Vigna |
| 2016 | ECIR | Compressing and Decoding Term Statistics Time Series. | Jinfeng Rao, Xing Niu, Jimmy Lin |
| 2016 | ICTIR | Total Recall: Blue Sky on Mars. | Charles L. A. Clarke, Gordon V. Cormack, Jimmy Lin, Adam Roegiest |
| 2016 | ICTIR | Rank-at-a-Time Query Processing. | Ahmed Elbagoury, Matt Crane, Jimmy Lin |
| 2016 | ICTIR | Retrievability in API-Based "Evaluation as a Service". | Jiaul H. Paik, Jimmy Lin |
| 2016 | ICTIR | Temporal Query Expansion Using a Continuous Hidden Markov Model. | Jinfeng Rao, Jimmy Lin |
| 2016 | NAACL | Pairwise Word Interaction Modeling with Deep Neural Networks for Semantic Similarity Measurement. | Hua He, Jimmy Lin |
| 2016 | SAC | Estimating topical volume in social media streams. | Praveen Bommannavar, Jimmy Lin, Anand Rajaraman |
| 2016 | SIGIR | Burst Detection in Social Media Streams for Tracking Interest Profiles in Real Time. | Cody Buntain, Jimmy Lin |
| 2016 | SIGIR | Interleaved Evaluation for Retrospective Summarization and Prospective Notification on Document Streams. | Xin Qian, Jimmy Lin, Adam Roegiest |
| 2016 | SIGIR | A Platform for Streaming Push Notifications to Mobile Assessors. | Adam Roegiest, Luchen Tan, Jimmy Lin, Charles L. A. Clarke |
| 2016 | SIGIR | Simple Dynamic Emission Strategies for Microblog Filtering. | Luchen Tan, Adam Roegiest, Charles L. A. Clarke, Jimmy Lin |
| 2016 | SIGIR | An Exploration of Evaluation Metrics for Mobile Push Notifications. | Luchen Tan, Adam Roegiest, Jimmy Lin, Charles L. A. Clarke |
| 2016 | SIGIR | Sampling Strategies and Active Learning for Volume Estimation. | Haotian Zhang, Jimmy Lin, Gordon V. Cormack, Mark D. Smucker |
| 2015 | ECIR | Reproducible Experiments on Lexical and Temporal Feedback for Tweet Search. | Jinfeng Rao, Jimmy Lin, Miles Efron |
| 2015 | EMNLP | Multi-Perspective Sentence Similarity Modeling with Convolutional Neural Networks. | Hua He, Kevin Gimpel, Jimmy Lin |
| 2015 | ICTIR | Building a Self-Contained Search Engine in the Browser. | Jimmy Lin |
| 2015 | ICTIR | Anytime Ranking for Impact-Ordered Indexes. | Jimmy Lin, Andrew Trotman |
| 2015 | ICTIR | The Feasibility of Brute Force Scans for Real-Time Tweet Search. | Yulu Wang, Jimmy Lin |
| 2015 | WWW | Scaling Down Distributed Infrastructure on Wimpy Machines for Personal Web Archiving. | Jimmy Lin |
| 2015 | SIGIR | SIGIR 2015 Workshop on Reproducibility, Inexplicability, and Generalizability of Results (RIGOR). | Jaime Arguello, Fernando Diaz, Jimmy Lin, Andrew Trotman |
| 2015 | SIGIR | Assessor Differences and User Preferences in Tweet Timeline Generation. | Yulu Wang, Garrick Sherman, Jimmy Lin, Miles Efron |
| 2014 | CSCW | Do recommendations matter?: news recommendation in real life. | Alan Said, Alejandro Bellogn, Jimmy Lin, Arjen P. de Vries |
| 2014 | ECIR | Column Stores as an IR Prototyping Tool. | Hannes Mhleisen, Thaer Samar, Jimmy Lin, Arjen P. de Vries |
| 2014 | ECIR | The Impact of Future Term Statistics in Real-Time Tweet Search. | Yulu Wang, Jimmy Lin |
| 2014 | EDBT | Optimization Techniques for "Scaling Down" Hadoop on Multi-Core, Shared-Memory Systems. | K. Ashwin Kumar, Jonathan Gluck, Amol Deshpande, Jimmy Lin |
| 2014 | WWW | Infrastructure support for evaluation as a service. | Jimmy Lin, Miles Efron |
| 2014 | WWW | Infrastructure for supporting exploration and discovery in web archives. | Jimmy Lin, Milad Gholami, Jinfeng Rao |
| 2014 | WWW | Information network or social network?: the structure of the twitter follow graph. | Seth A. Myers, Aneesh Sharma, Pankaj Gupta, Jimmy Lin |
| 2014 | WWW | Learning to efficiently rank on big data. | Lidan Wang, Jimmy Lin, Donald Metzler, Jiawei Han |
| 2014 | SIGIR | Temporal feedback for tweet search with non-parametric density estimation. | Miles Efron, Jimmy Lin, Jiyin He, Arjen P. de Vries |
| 2014 | SIGIR | Old dogs are great at new tricks: column stores for ir prototyping. | Hannes Mhleisen, Thaer Samar, Jimmy Lin, Arjen P. de Vries |
| 2014 | SIGIR | On run diversity in Evaluation as a Service. | Ellen M. Voorhees, Jimmy Lin, Miles Efron |
| 2013 | ACL | Mr. MIRA: Open-Source Large-Margin Structured Learning on MapReduce. | Vladimir Eidelman, Ke Wu, Ferhan Tre, Philip Resnik, Jimmy Lin |
| 2013 | CIKM | A month in the life of a production news recommender system. | Alan Said, Jimmy Lin, Alejandro Bellogn, Arjen P. de Vries |
| 2013 | ECIR | Training Efficient Tree-Based Models for Document Ranking. | Sebastian Bruch, Jimmy Lin |
| 2013 | ICWSM | Visualizing the "Pulse" of World Cities on Twitter. | Miguel Rios, Jimmy Lin |
| 2013 | KDD | Dynamic memory allocation policies for postings in real-time Twitter search. | Sebastian Bruch, Jimmy Lin, Michael Busch |
| 2013 | NAACL | Massively Parallel Suffix Array Queries and On-Demand Phrase Extraction for Statistical Machine Translation Using GPUs. | Hua He, Jimmy Lin, Adam Lopez |
| 2013 | WWW | WTF: the who to follow service at Twitter. | Pankaj Gupta, Ashish Goel, Jimmy Lin, Aneesh Sharma, Dong Wang, Reza Zadeh |
| 2013 | SIGIR | Effectiveness/efficiency tradeoffs for candidate generation in multi-stage retrieval architectures. | Sebastian Bruch, Jimmy Lin |
| 2013 | SIGIR | Flat vs. hierarchical phrase-based translation models for cross-language information retrieval. | Ferhan Tre, Jimmy Lin |
| 2012 | CIKM | Fast candidate generation for two-phase document ranking: postings list intersection with bloom filters. | Sebastian Bruch, Jimmy Lin |
| 2012 | COLING | Combining Statistical Translation Techniques for Cross-Language Information Retrieval. | Ferhan Tre, Jimmy Lin, Douglas W. Oard |
| 2012 | ICDE | Earlybird: Real-Time Search at Twitter. | Michael Busch, Krishna Gade, Brian Larson, Patrick Lok, Samuel Luckenbill, Jimmy Lin |
| 2012 | ICWSM | A Study of "Churn" in Tweets and Real-Time Search Queries. | Jimmy Lin, Gilad Mishne |
| 2012 | ICWSM | Evaluating Real-Time Search over Tweets. | Dean McCullough, Jimmy Lin, Craig Macdonald, Iadh Ounis, Richard McCreadie |
| 2012 | NAACL | Why Not Grab a Free Lunch? Mining Large Corpora for Parallel Sentences to Improve Translation Modeling. | Ferhan Tre, Jimmy Lin |
| 2012 | SIGIR | On building a reusable Twitter corpus. | Richard McCreadie, Ian Soboroff, Jimmy Lin, Craig Macdonald, Iadh Ounis, Dean McCullough |
| 2012 | SIGIR | Twanchor text: a preliminary study of the value of tweets as anchor text. | Gilad Mishne, Jimmy Lin |
| 2012 | SIGIR | Looking inside the box: context-sensitive translation for cross-language information retrieval. | Ferhan Tre, Jimmy Lin, Douglas W. Oard |
| 2011 | CIKM | When close enough is good enough: approximate positional indexes for efficient ranked retrieval. | Tamer Elsayed, Jimmy Lin, Donald Metzler |
| 2011 | CLOUD | Automatic management of partitioned, replicated search services. | Florian Leibert, Jake Mannix, Jimmy Lin, Babak Hamadani |
| 2011 | KDD | Smoothing techniques for adaptive online language models: topic tracking in tweet streams. | Jimmy Lin, Rion Snow, William Morgan |
| 2011 | SIGIR | Pseudo test collections for learning web search ranking functions. | Sebastian Bruch, Donald Metzler, Tamer Elsayed, Jimmy Lin |
| 2011 | SIGIR | Cross-corpus relevance projection. | Sebastian Bruch, Donald Metzler, Jimmy Lin |
| 2011 | SIGIR | No free lunch: brute force vs. locality-sensitive hashing for cross-lingual pairwise similarity. | Ferhan Tre, Tamer Elsayed, Jimmy Lin |
| 2011 | SIGIR | A cascade ranking model for efficient ranked retrieval. | Lidan Wang, Jimmy Lin, Donald Metzler |
| 2010 | CIKM | Ranking under temporal constraints. | Lidan Wang, Donald Metzler, Jimmy Lin |
| 2010 | CloudCom | Scaling Populations of a Genetic Algorithm for Job Shop Scheduling Problems Using MapReduce. | Di-Wei Huang, Jimmy Lin |
| 2010 | NAACL | Data-Intensive Text Processing with MapReduce. | Jimmy Lin, Chris Dyer |
| 2010 | NAACL | Putting the User in the Loop: Interactive Maximal Marginal Relevance for Query-Focused Summarization. | Jimmy Lin, Nitin Madnani, Bonnie J. Dorr |
| 2010 | SIGIR | Learning to efficiently rank. | Lidan Wang, Jimmy Lin, Donald Metzler |
| 2009 | ICWSM | You Are Where You Edit: Locating Wikipedia Contributors through Edit Histories. | Michael D. Lieberman, Jimmy Lin |
| 2009 | NAACL | Data Intensive Text Processing with MapReduce. | Jimmy Lin, Chris Dyer |
| 2009 | SIGIR | Brute force and indexed approaches to pairwise document similarity comparisons with MapReduce. | Jimmy Lin |
| 2009 | SIGIR | The Curse of Zipf and Limits to Parallelization: An Look at the Stragglers Problem in MapReduce. | Jimmy Lin |
| 2008 | ACL | Pairwise Document Similarity in Large Collections with MapReduce. | Tamer Elsayed, Jimmy Lin, Douglas W. Oard |
| 2008 | EMNLP | Scalable Language Processing Algorithms for the Masses: A Case Study in Computing Word Co-occurrence Matrices with MapReduce. | Jimmy Lin |
| 2008 | SIGIR | How do users find things with PubMed?: towards automatic utility evaluation with user simulations. | Jimmy Lin, Mark D. Smucker |
| 2007 | ACL | Different Structures for Evaluating Answers to Complex Questions: Pyramids Won't Topple, and Neither Will Human Assessors. | Hoa Trang Dang, Jimmy Lin |
| 2007 | AMIA | Semantic Clustering of Answers to Clinical Questions. | Jimmy Lin, Dina Demner-Fushman |
| 2007 | NAACL | Is Question Answering Better than Information Retrieval? Towards a Task-Based Evaluation Framework for Question Series. | Jimmy Lin |
| 2007 | SIGIR | Deconstructing nuggets: the stability and reliability of complex question answering evaluation. | Jimmy Lin, Pengyi Zhang |
| 2006 | ACL | Answer Extraction, Semantic Clustering, and Extractive Summarization for Clinical Question Answering. | Dina Demner-Fushman, Jimmy Lin |
| 2006 | ACL | The Role of Information Retrieval in Answering Complex Questions. | Jimmy Lin |
| 2006 | ACL | Leveraging Reusability: Cost-Effective Lexical Acquisition for Large-Scale Ontology Translation. | G. Craig Murray, Bonnie J. Dorr, Jimmy Lin, Jan Hajic, Pavel Pecina |
| 2006 | AMIA | Evaluation of PICO as a Knowledge Representation for Clinical Questions. | Xiaoli Huang, Jimmy Lin, Dina Demner-Fushman |
| 2006 | EAMT | Leveraging Recurrent Phrase Structure in Large-scale Ontology Translation. | G. Craig Murray, Bonnie J. Dorr, Jimmy Lin, Jan Hajic, Pavel Pecina |
| 2006 | NAACL | Will Pyramids Built of Nuggets Topple Over?. | Jimmy Lin, Dina Demner-Fushman |
| 2006 | SIGIR | The role of knowledge in conceptual retrieval: a study in the domain of clinical medicine. | Jimmy Lin, Dina Demner-Fushman |
| 2006 | SIGIR | Exploring the limits of single-iteration clarification dialogs. | Jimmy Lin, Philip Fei Wu, Dina Demner-Fushman, Eileen G. Abels |
| 2006 | SIGIR | Action modeling: language models that predict query behavior. | G. Craig Murray, Jimmy Lin, Abdur Chowdhury |
| 2005 | ACL | Evaluating Summaries and Answers: Two Sides of the Same Coin? | Jimmy Lin, Dina Demner-Fushman |
| 2005 | AMIA | "Bag of Words" is not enough for Strength of Evidence Classification. | Jimmy Lin, Dina Demner-Fushman |
| 2005 | NAACL | Automatically Evaluating Answers to Definition Questions. | Jimmy Lin, Dina Demner-Fushman |
| 2005 | SIGIR | Evaluation of resources for question answering evaluation. | Jimmy Lin |
| 2005 | SIGIR | Assessing the term independence assumption in blind relevance feedback. | Jimmy Lin, G. Craig Murray |
| 2004 | NAACL | Answering Definition Questions Using Multiple Knowledge Sources. | Wesley Hildebrandt, Boris Katz, Jimmy Lin |
| 2004 | NAACL | A Computational Framework for Non-Lexicalist Semantics. | Jimmy Lin |
| 2003 | CHI | The role of context in question answering systems. | Jimmy Lin, Dennis Quan, Vineet Sinha, Karun Bakshi, David Huynh, Boris Katz, David R. Karger |
| 2003 | CIKM | Question answering from the web using knowledge annotation and knowledge mining techniques. | Jimmy Lin, Boris Katz |
| 2003 | Interact | What Makes a Good Answer? The Role of Context in Question Answering. | Jimmy Lin, Dennis Quan, Vineet Sinha, Karun Bakshi, David Huynh, Boris Katz, David R. Karger |
| 2003 | IUI | Sticky notes for the semantic web. | David R. Karger, Boris Katz, Jimmy Lin, Dennis Quan |
| 2003 | SIGIR | Quantitative evaluation of passage retrieval algorithms for question answering. | Stefanie Tellex, Boris Katz, Jimmy Lin, Aaron Fernandes, Gregory Marton |
| 2002 | CoopIS | Natural Language Annotations for the Semantic Web. | Boris Katz, Jimmy Lin, Dennis Quan |
| 2002 | LREC | The Web as a Resource for Question Answering: Perspectives and Challenges. | Jimmy Lin |
| 2002 | NLDB | Omnibase: Uniform Access to Heterogeneous Data for Question Answering. | Boris Katz, Sue Felshin, Deniz Yuret, Ali Ibrahim, Jimmy Lin, Gregory Marton, Alton Jerome McFarland, Baris Temelkuran |
| 2002 | SIGIR | Web question answering: is more always better?. | Susan T. Dumais, Michele Banko, Eric Brill, Jimmy Lin, Andrew Y. Ng |
| 2001 | ACL | Gathering Knowledge for a Question Answering System from Heterogeneous Information Sources. | Boris Katz, Jimmy Lin, Sue Felshin |