| 2026 | ACL | e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings. | Haonan Chen, Sicheng Gao, Radu Timofte, Tetsuya Sakai, Zhicheng Dou |
| 2026 | WWW | OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAG. | Fengran Mo, Zhan Su, Yuchen Hui, Jinghan Zhang, Jia Ao Sun, Zheyuan Liu, Chao Zhang, Tetsuya Sakai, Jian-Yun Nie |
| 2026 | WSDM | Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Simulation. | Nuo Chen, Hanpei Fang, Jiqun Liu, Wilson Wei, Tetsuya Sakai, Xiao-Ming Wu |
| 2026 | WSDM | Diversification as Risk Minimization. | Rikiya Takehi, Fernando Diaz, Tetsuya Sakai |
| 2025 | CIKM | Open-Source LLM-based Relevance Assessment vs. Highly Reliable Manual Relevance Assessment: A Case Study. | Tetsuya Sakai, Khant Myoe Rain, Rikiya Takehi, Sijie Tao, Young-In Song |
| 2025 | ICPRAM | Reconstruction of 3D Brain Structures from Clinical 2D MRI Data. | Rui Shi, Tsukasa Koike, Tetsuro Sekine, Akio Morita, Tetsuya Sakai |
| 2025 | NAACL | CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmented Generation. | Yiruo Cheng, Kelong Mao, Ziliang Zhao, Guanting Dong, Hongjin Qian, Yongkang Wu, Tetsuya Sakai, Ji-Rong Wen, Zhicheng Dou |
| 2025 | SIGIR | My System Is As Effective As Yours: Reproducibility, Sustainability, and More. | Tetsuya Sakai |
| 2025 | SIGIR | LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help? | Rikiya Takehi, Ellen M. Voorhees, Tetsuya Sakai, Ian Soboroff |
| 2024 | EMNLP | ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense Retrieval. | Kelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo, Zheng Liu, Tetsuya Sakai, Zhicheng Dou |
| 2024 | EMNLP | ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models. | Yuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu, Cheng Yang, Chufan Shi, Xinyu Zhu, Zihao Lin, Hanwen Wan, Yujiu Yang, Tetsuya Sakai, Tian Feng, Hayato Yamana |
| 2024 | WSDM | ONCE: Boosting Content-based Recommendation with Both Open- and Closed-source Large Language Models. | Qijiong Liu, Nuo Chen, Tetsuya Sakai, Xiao-Ming Wu |
| 2023 | CVPR | MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model. | Yatai Ji, Junjie Wang, Yuan Gong, Lin Zhang, Yanru Zhu, Hongfa Wang, Jiaxing Zhang, Tetsuya Sakai, Yujiu Yang |
| 2023 | ICTIR | Evaluating Parrots and Sociopathic Liars (keynote). | Tetsuya Sakai |
| 2023 | WWW | A Reference-Dependent Model for Web Search Evaluation: Understanding and Measuring the Experience of Boundedly Rational Users. | Nuo Chen, Jiqun Liu, Tetsuya Sakai |
| 2023 | SIGIR | Practice and Challenges in Building a Business-oriented Search Engine Quality Metric. | Nuo Chen, Donghyun Park, Hyungae Park, Kijun Choi, Tetsuya Sakai, Jinyoung Kim |
| 2022 | COLING | LayerConnect: Hypernetwork-Assisted Inter-Layer Connector to Enhance Parameter Efficiency. | Haoxiang Shi, Rongsheng Zhang, Jiaan Wang, Cen Wang, Yinhe Zheng, Tetsuya Sakai |
| 2022 | CVPR | AxIoU: An Axiomatically Justified Measure for Video Moment Retrieval. | Riku Togashi, Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkil, Tetsuya Sakai |
| 2022 | EMNLP | Zero-Shot Learners for Natural Language Understanding via a Unified Multiple Choice Perspective. | Ping Yang, Junjie Wang, Ruyi Gan, Xinyu Zhu, Lin Zhang, Ziwei Wu, Xinyu Gao, Jiaxing Zhang, Tetsuya Sakai |
| 2022 | ICTIR | Do Extractive Summarization Algorithms Amplify Lexical Bias in News Articles? | Rei Shimizu, Sumio Fujita, Tetsuya Sakai |
| 2022 | LREC | Evaluating the Effects of Embedding with Speaker Identity Information in Dialogue Summarization. | Yuji Naraki, Tetsuya Sakai, Yoshihiko Hayashi |
| 2022 | RAID | Understanding the Behavior Transparency of Voice Assistant Applications Using the ChatterBox Framework. | Atsuko Natatsuka, Ryo Iijima, Takuya Watanabe, Mitsuaki Akiyama, Tetsuya Sakai, Tatsuya Mori |
| 2022 | SIGIR | Constructing Better Evaluation Metrics by Incorporating the Anchoring Effect into the User Model. | Nuo Chen, Fan Zhang, Tetsuya Sakai |
| 2021 | ACL | Evaluating Evaluation Measures for Ordinal Classification and Ordinal Quantification. | Tetsuya Sakai |
| 2021 | CIKM | Incorporating Query Reformulating Behavior into Web Search Evaluation. | Jia Chen, Yiqun Liu, Jiaxin Mao, Fan Zhang, Tetsuya Sakai, Weizhi Ma, Min Zhang, Shaoping Ma |
| 2021 | CIKM | Evaluating Relevance Judgments with Pairwise Discriminative Power. | Zhumin Chu, Jiaxin Mao, Fan Zhang, Yiqun Liu, Tetsuya Sakai, Min Zhang, Shaoping Ma |
| 2021 | CIKM | A Closer Look at Evaluation Measures for Ordinal Quantification. | Tetsuya Sakai |
| 2021 | ECIR | How Do Users Revise Zero-Hit Product Search Queries? | Yuki Amemiya, Tomohiro Manabe, Sumio Fujita, Tetsuya Sakai |
| 2021 | ECIR | On the Instability of Diminishing Return IR Measures. | Tetsuya Sakai |
| 2021 | EMNLP | MIRTT: Learning Multimodal Interaction Representations from Trilinear Transformers for Visual Question Answering. | Junjie Wang, Yatai Ji, Jiaqi Sun, Yujiu Yang, Tetsuya Sakai |
| 2021 | ICTIR | A Fast and Exact Randomisation Test for Comparing Two Systems with Paired Data. | Rikiya Suzuki, Tetsuya Sakai |
| 2021 | SMC | A Simple and Effective Usage of Self-supervised Contrastive Learning for Text Clustering. | Haoxiang Shi, Cen Wang, Tetsuya Sakai |
| 2021 | SIGIR | On the Two-Sample Randomisation Test for IR Evaluation. | Tetsuya Sakai |
| 2021 | SIGIR | WWW3E8: 259, 000 Relevance Labels for Studying the Effect of Document Presentation Order for Relevance Assessors. | Tetsuya Sakai, Sijie Tao, Zhaohao Zeng |
| 2021 | SIGIR | Scalable Personalised Item Ranking through Parametric Density Estimation. | Riku Togashi, Masahiro Kato, Mayu Otani, Tetsuya Sakai, Shin'ichi Satoh |
| 2020 | IJCNLP | A Siamese CNN Architecture for Learning Chinese Sentence Similarity. | Haoxiang Shi, Cen Wang, Tetsuya Sakai |
| 2020 | SIGIR | How to Measure the Reproducibility of System-oriented IR Experiments. | Timo Breuer, Nicola Ferro, Norbert Fuhr, Maria Maistro, Tetsuya Sakai, Philipp Schaer, Ian Soboroff |
| 2020 | SIGIR | Good Evaluation Measures based on Document Preferences. | Tetsuya Sakai, Zhaohao Zeng |
| 2020 | SIGIR | Visual Intents vs. Clicks, Likes, and Purchases in E-commerce. | Riku Togashi, Tetsuya Sakai |
| 2019 | CCS | Poster: A First Look at the Privacy Risks of Voice Assistant Apps. | Atsuko Natatsuka, Ryo Iijima, Takuya Watanabe, Mitsuaki Akiyama, Tetsuya Sakai, Tatsuya Mori |
| 2019 | ECIR | CENTRE@CLEF 2019. | Nicola Ferro, Norbert Fuhr, Maria Maistro, Tetsuya Sakai, Ian Soboroff |
| 2019 | ICTIR | Generalising Kendall's Tau for Noisy and Incomplete Preference Judgements. | Riku Togashi, Tetsuya Sakai |
| 2019 | SMC | System Evaluation of Ternary Error-Correcting Output Codes for Multiclass Classification Problems. | Shigeichi Hirasawa, Gendo Kumoi, Hideki Yagi, Manabu Kobayashi, Masayuki Goto, Tetsuya Sakai, Hiroshige Inazumi |
| 2019 | SIGIR | The SIGIR 2019 Open-Source IR Replicability Challenge (OSIRRC 2019). | Ryan Clancy, Nicola Ferro, Claudia Hauff, Jimmy Lin, Tetsuya Sakai, Ze Zhong Wu |
| 2019 | SIGIR | Overview of the 2019 Open-Source IR Replicability Challenge (OSIRRC 2019). | Ryan Clancy, Nicola Ferro, Claudia Hauff, Jimmy Lin, Tetsuya Sakai, Ze Zhong Wu |
| 2019 | SIGIR | Which Diversity Evaluation Measures Are "Good"? | Tetsuya Sakai, Zhaohao Zeng |
| 2019 | SIGIR | BM25 Pseudo Relevance Feedback Using Anserini at Waseda University. | Zhaohao Zeng, Tetsuya Sakai |
| 2019 | UIST | Voice Input Interface Failures and Frustration: Developer and User Perspectives. | Shiyoh Goetsu, Tetsuya Sakai |
| 2019 | WSDM | Conducting Laboratory Experiments Properly with Statistical Tools: An Easy Hands-On Tutorial. | Tetsuya Sakai |
| 2019 | WSDM | Attitude Detection for One-Round Conversation: Jointly Extracting Target-Polarity Pairs. | Zhaohao Zeng, Ruihua Song, Pingping Lin, Tetsuya Sakai |
| 2018 | ICDM | Why You Should Listen to This Song: Reason Generation for Explainable Recommendation. | Guoshuai Zhao, Hao Fu, Ruihua Song, Tetsuya Sakai, Xing Xie, Xueming Qian |
| 2018 | ICTIR | Topic Set Size Design for Paired and Unpaired Data. | Tetsuya Sakai |
| 2018 | ICTIR | Classifying Community QA Questions That Contain an Image. | Kenta Tamaki, Riku Togashi, Sosuke Kato, Sumio Fujita, Hideyuki Maeda, Tetsuya Sakai |
| 2018 | SIGIR | Comparing Two Binned Probability Distributions for Information Access Evaluation. | Tetsuya Sakai |
| 2018 | SIGIR | Conducting Laboratory Experiments Properly with Statistical Tools: An Easy Hands-on Tutorial. | Tetsuya Sakai |
| 2017 | CHIIR | Investigating Users' Time Perception during Web Search. | Cheng Luo, Xue Li, Yiqun Liu, Tetsuya Sakai, Fan Zhang, Min Zhang, Shaoping Ma |
| 2017 | CIKM | Ranking Rich Mobile Verticals based on Clicks and Abandonment. | Mami Kawasaki, Inho Kang, Tetsuya Sakai |
| 2017 | ICTIR | Mobile Vertical Ranking based on Preference Graphs. | Yuta Kadotami, Yasuaki Yoshida, Sumio Fujita, Tetsuya Sakai |
| 2017 | SIGIR | LSTM vs. BM25 for Open-domain QA: A Hands-on Comparison of Effectiveness and Efficiency. | Sosuke Kato, Riku Togashi, Hideyuki Maeda, Sumio Fujita, Tetsuya Sakai |
| 2017 | SIGIR | Evaluating Mobile Search with Height-Biased Gain. | Cheng Luo, Yiqun Liu, Tetsuya Sakai, Fan Zhang, Min Zhang, Shaoping Ma |
| 2017 | SIGIR | The Probability that Your Hypothesis Is Correct, Credible Intervals, and Effect Sizes for IR Evaluation. | Tetsuya Sakai |
| 2017 | WSDM | Does Document Relevance Affect the Searcher's Perception of Time? | Cheng Luo, Yiqun Liu, Tetsuya Sakai, Ke Zhou, Fan Zhang, Xue Li, Shaoping Ma |
| 2016 | ICTIR | Topic Set Size Design and Power Analysis in Practice. | Tetsuya Sakai |
| 2016 | ICTIR | Simple and Effective Approach to Score Standardisation. | Tetsuya Sakai |
| 2016 | SIGIR | Statistical Significance, Power, and Sample Sizes: A Systematic Review of SIGIR and TOIS, 2006-2015. | Tetsuya Sakai |
| 2016 | SIGIR | Two Sample T-tests for IR Evaluation: Student or Welch? | Tetsuya Sakai |
| 2016 | SIGIR | Evaluating Search Result Diversity using Intent Hierarchies. | Xiaojie Wang, Zhicheng Dou, Tetsuya Sakai, Ji-Rong Wen |
| 2015 | CIKM | ECol 2015: First international workshop on the Evaluation on Collaborative Information Seeking and Retrieval. | Leif Azzopardi, Jeremy Pickens, Tetsuya Sakai, Laure Soulier, Lynda Tamine |
| 2015 | CIKM | Search Result Diversification Based on Hierarchical Intents. | Sha Hu, Zhicheng Dou, Xiaojie Wang, Tetsuya Sakai, Ji-Rong Wen |
| 2015 | SOUPS | Understanding the Inconsistencies between Text Descriptions and the Use of Privacy-sensitive Resources of Mobile Apps. | Takuya Watanabe, Mitsuaki Akiyama, Tetsuya Sakai, Tatsuya Mori |
| 2014 | CIKM | Designing Test Collections for Comparing Many Systems. | Tetsuya Sakai |
| 2013 | CIKM | Dynamic query intent mining from a search log stream. | Ya-nan Qian, Tetsuya Sakai, Junting Ye, Qinghua Zheng, Cong Li |
| 2013 | CIKM | On the reliability and intuitiveness of aggregated search metrics. | Ke Zhou, Mounia Lalmas, Tetsuya Sakai, Ronan Cummins, Joemon M. Jose |
| 2013 | SIGIR | Exploring semi-automatic nugget extraction for Japanese one click access evaluation. | Matthew Ekstrand-Abueg, Virgil Pavlu, Makoto P. Kato, Tetsuya Sakai, Takehiro Yamamoto, Mayu Iwata |
| 2013 | SIGIR | Report from the NTCIR-10 1CLICK-2 Japanese subtask: baselines, upperbounds and evaluation robustness. | Makoto P. Kato, Tetsuya Sakai, Takehiro Yamamoto, Mayu Iwata |
| 2013 | SIGIR | Time-aware structured query suggestion. | Taiki Miyanishi, Tetsuya Sakai |
| 2013 | SIGIR | Summaries, ranked retrieval and sessions: a unified framework for information access evaluation. | Tetsuya Sakai, Zhicheng Dou |
| 2013 | SIGIR | The impact of intent selection on diversified search evaluation. | Tetsuya Sakai, Zhicheng Dou, Charles L. A. Clarke |
| 2013 | SIGIR | Summary of the NTCIR-10 INTENT-2 task: subtopic mining and search result diversification. | Tetsuya Sakai, Zhicheng Dou, Takehiro Yamamoto, Yiqun Liu, Min Zhang, Makoto P. Kato, Ruihua Song, Mayu Iwata |
| 2012 | CIKM | The wisdom of advertisers: mining subgoals via query clustering. | Takehiro Yamamoto, Tetsuya Sakai, Mayu Iwata, Chen Yu, Ji-Rong Wen, Katsumi Tanaka |
| 2012 | WWW | Structured query suggestion for specialization and parallel movement: effect on search behaviors. | Makoto P. Kato, Tetsuya Sakai, Katsumi Tanaka |
| 2012 | WWW | Evaluation with informational and navigational intents. | Tetsuya Sakai |
| 2012 | SIGIR | AspecTiles: tile-based visualization of diversified web search results. | Mayu Iwata, Tetsuya Sakai, Takehiro Yamamoto, Yu Chen, Yi Liu, Ji-Rong Wen, Shojiro Nishio |
| 2012 | SIGIR | New assessment criteria for query suggestion. | Zhongrui Ma, Yu Chen, Ruihua Song, Tetsuya Sakai, Jiaheng Lu, Ji-Rong Wen |
| 2012 | SIGIR | Towards zero-click mobile IR evaluation: knowing what and knowing when. | Tetsuya Sakai |
| 2011 | ACL | Query Snowball: A Co-occurrence-based Approach to Multi-document Summarization for Question Answering. | Hajime Morita, Tetsuya Sakai, Manabu Okumura |
| 2011 | CIKM | Click the search button and be happy: evaluating direct and immediate information access. | Tetsuya Sakai, Makoto P. Kato, Young-In Song |
| 2011 | SIGIR | Evaluating diversified search results using per-intent graded relevance. | Tetsuya Sakai, Ruihua Song |
| 2011 | WSDM | Using graded-relevance metrics for evaluating community QA answer selection. | Tetsuya Sakai, Daisuke Ishikawa, Noriko Kando, Yohei Seki, Kazuko Kuriyama, Chin-Yew Lin |
| 2009 | SIGIR | Serendipitous search via wikipedia: a query log analysis. | Tetsuya Sakai, Kenichi Nogami |
| 2008 | CIKM | Comparing metrics across TREC and NTCIR: the robustness to system bias. | Tetsuya Sakai |
| 2008 | SIGIR | Comparing metrics across TREC and NTCIR: : the robustness to pool depth bias. | Tetsuya Sakai |
| 2008 | SIGIR | Precision-at-ten considered redundant. | William Webber, Alistair Moffat, Justin Zobel, Tetsuya Sakai |
| 2007 | SIGIR | Alternatives to Bpref. | Tetsuya Sakai |
| 2006 | SIGIR | Evaluating evaluation metrics based on the bootstrap. | Tetsuya Sakai |
| 2006 | SIGIR | Give me just one highly relevant document: P-measure. | Tetsuya Sakai |
| 2004 | SIGIR | The effect of back-formulating questions in question answering evaluation. | Tetsuya Sakai, Yoshimi Saito, Yumi Ichimura, Tomoharu Kokubu, Makoto Koyama |
| 2003 | SIGIR | Average gain ratio: a simple retrieval performance measure for evaluation with multiple relevance levels. | Tetsuya Sakai |
| 2003 | SIGIR | Evaluating retrieval performance for Japanese question answering: what are best passages? | Tetsuya Sakai, Tomoharu Kokubu |
| 2002 | SMC | The use of external text data in cross-language information retrieval based on machine translation. | Tetsuya Sakai |
| 2002 | SMC | Generating transliteration rules for cross-language information retrieval from machine translation dictionaries. | Tetsuya Sakai, Akira Kumano, Toshihiko Manabe |
| 2002 | SIGIR | Relative and absolute term selection criteria: a comparative study for English and Japanese IR. | Tetsuya Sakai, Stephen E. Robertson |
| 2001 | SIGIR | Generic Summaries for Indexing in Information Retrieval. | Tetsuya Sakai, Karen Sparck Jones |
| 2001 | SIGIR | Flexible Pseudo-Relevance Feedback Using Optimization Tables. | Tetsuya Sakai, Stephen E. Robertson |
| 1999 | SIGIR | A Comparison of Query Translation Methods for English-Japanese Cross-Language Information Retrieval (poster abstract). | Gareth J. F. Jones, Tetsuya Sakai, Nigel Collier, Akira Kumano, Kazuo Sumita |
| 1998 | SIGIR | Experiments in Japanese Text Retrieval and Routing Using the NEAT System. | Gareth J. F. Jones, Tetsuya Sakai, Masahiro Kajiura, Kazuo Sumita |
| 1998 | SIGIR | Lessons from BMIR-J2: A Test Collection for Japanese IR Systems. | Tsuyoshi Kitani, Yasushi Ogawa, Tetsuya Ishikawa, Haruo Kimoto, Ikuo Keshi, Jun Toyoura, Toshikazu Fukushima, Kunio Matsui, Yoshihiro Ueda, Tetsuya Sakai, Takenobu Tokunaga, Hiroshi Tsuruoka, Hidekazu Nakawatase, Teru Agata |