| 2026 | ACL | Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations. | Christabel Acquaye, Yi Ting Huang, Marine Carpuat, Rachel Rudinger |
| 2026 | ACL | Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers. | Nishant Balepur, Atrey Desai, Rachel Rudinger |
| 2026 | ACL | BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks. | Nishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Vipul Gupta, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan Lee Boyd-Graber |
| 2026 | ACL | Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms. | Dayeon Ki, Yu Hou, Rachel Rudinger, Hal Daum III, Marine Carpuat, Fumeng Yang |
| 2026 | ACL | Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility. | Shramay Palta, Peter Rankel, Sarah Wiegreffe, Rachel Rudinger |
| 2025 | ACL | On the Mutual Influence of Gender and Occupation in LLM Representations. | Haozhe An, Connor Baumler, Abhilasha Sancheti, Rachel Rudinger |
| 2025 | ACL | Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas. | Nishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng, Rachel Rudinger, Jordan Lee Boyd-Graber |
| 2025 | ACL | Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above. | Nishant Balepur, Rachel Rudinger, Jordan Lee Boyd-Graber |
| 2025 | ACL | Multiple LLM Agents Debate for Equitable Cultural Alignment. | Dayeon Ki, Rachel Rudinger, Tianyi Zhou, Marine Carpuat |
| 2025 | ACL | Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs. | Rupak Sarkar, Neha Srikanth, Taylor Pellegrin, Rachel Rudinger, Claire Bonial, Philip Resnik |
| 2025 | ACL | No Questions are Stupid, but some are Poorly Posed: Understanding Poorly-Posed Information-Seeking Questions. | Neha Srikanth, Rachel Rudinger, Jordan Lee Boyd-Graber |
| 2025 | EMNLP | A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users. | Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, Jordan Lee Boyd-Graber |
| 2025 | EMNLP | 'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ? | Huy Nghiem, Phuong-Anh Nguyen-Le, John Prindle, Rachel Rudinger, Hal Daum III |
| 2025 | ICLR | Natural Language Inference Improves Compositionality in Vision-Language Models. | Paola Cascante-Bonilla, Yu Hou, Yang Trista Cao, Hal Daum III, Rachel Rudinger |
| 2025 | IJCNLP | Speaking the Right Language: The Impact of Expertise (Mis)Alignment in User-AI Interactions. | Shramay Palta, Nirupama Chandrasekaran, Rachel Rudinger, Scott Counts |
| 2025 | NAACL | Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer? | Nishant Balepur, Feng Gu, Abhilasha Ravichander, Shi Feng, Jordan Lee Boyd-Graber, Rachel Rudinger |
| 2025 | NAACL | Language Models Predict Empathy Gaps Between Social In-groups and Out-groups. | Yu Hou, Hal Daum III, Rachel Rudinger |
| 2025 | NAACL | NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals. | Neha Srikanth, Rachel Rudinger |
| 2024 | ACL | Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender? | Haozhe An, Christabel Acquaye, Colin Wang, Zongxia Li, Rachel Rudinger |
| 2024 | ACL | It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning. | Nishant Balepur, Shramay Palta, Rachel Rudinger |
| 2024 | ACL | Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question? | Nishant Balepur, Abhilasha Ravichander, Rachel Rudinger |
| 2024 | CogSci | Assessing Common Ground through Language-based Cultural Consensus in Humans and Large Language Models. | Sophie Domanski, Rachel Rudinger, Marine Carpuat, Patrick Shafto, Yi Ting Huang |
| 2024 | EMNLP | Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the US. | Christabel Acquaye, Haozhe An, Rachel Rudinger |
| 2024 | EMNLP | Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning. | Shramay Palta, Nishant Balepur, Peter Rankel, Sarah Wiegreffe, Marine Carpuat, Rachel Rudinger |
| 2024 | EMNLP | On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models. | Abhilasha Sancheti, Haozhe An, Rachel Rudinger |
| 2024 | NAACL | Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering. | Neha Srikanth, Rupak Sarkar, Heran Mane, Elizabeth Aparicio, Quynh C. Nguyen, Rachel Rudinger, Jordan L. Boyd-Graber |
| 2023 | ACL | Nichelle and Nancy: The Influence of Demographic Attributes and Tokenization Length on First Name Biases. | Haozhe An, Rachel Rudinger |
| 2023 | ACL | FORK: A Bite-Sized Test Set for Probing Culinary Cultural Biases in Commonsense Reasoning Models. | Shramay Palta, Rachel Rudinger |
| 2023 | EACL | SODAPOP: Open-Ended Discovery of Social Biases in Social Commonsense Reasoning Models. | Haozhe An, Zongxia Li, Jieyu Zhao, Rachel Rudinger |
| 2023 | EMNLP | What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and Prohibitions. | Abhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, Rachel Rudinger |
| 2022 | AAAI | Entailment Relation Aware Paraphrase Generation. | Abhilasha Sancheti, Balaji Vasan Srinivasan, Rachel Rudinger |
| 2022 | EMNLP | Agent-Specific Deontic Modality Detection in Legal Language. | Abhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, Rachel Rudinger |
| 2022 | NAACL | Recognition of They/Them as Singular Personal Pronouns in Coreference Resolution. | Connor Baumler, Rachel Rudinger |
| 2022 | NAACL | Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models. | Yang Trista Cao, Anna Sotnikova, Hal Daum III, Rachel Rudinger, Linda Zou |
| 2022 | NAACL | Partial-input baselines show that NLI models can ignore context, but they don't. | Neha Srikanth, Rachel Rudinger |
| 2021 | AAAI | Learning to Rationalize for Nonmonotonic Reasoning with Distant Supervision. | Faeze Brahman, Vered Shwartz, Rachel Rudinger, Yejin Choi |
| 2021 | ACL | MedNLI Is Not Immune: Natural Language Inference Artifacts in the Clinical Domain. | Christine Herlihy, Rachel Rudinger |
| 2021 | ACL | Analyzing Stereotypes in Generative Text Inference Tasks. | Anna Sotnikova, Yang Trista Cao, Hal Daum III, Rachel Rudinger |
| 2020 | EMNLP | Thinking Like a Skeptic: Defeasible Inference in Natural Language. | Rachel Rudinger, Vered Shwartz, Jena D. Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A. Smith, Yejin Choi |
| 2020 | EMNLP | "You are grounded!": Latent Name Artifacts in Pre-trained Language Models. | Vered Shwartz, Rachel Rudinger, Oyvind Tafjord |
| 2020 | EMNLP | Causal Inference of Script Knowledge. | Noah Weber, Rachel Rudinger, Benjamin Van Durme |
| 2020 | LREC | The Universal Decompositional Semantics Dataset and Decomp Toolkit. | Aaron Steven White, Elias Stengel-Eskin, Siddharth Vashishtha, Venkata Subrahmanyan Govindarajan, Dee Ann Reisinger, Tim Vieira, Keisuke Sakaguchi, Sheng Zhang, Francis Ferraro, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
| 2019 | AAAI | PARABANK: Monolingual Bitext Generation and Sentential Paraphrasing via Lexically-Constrained Neural Machine Translation. | J. Edward Hu, Rachel Rudinger, Matt Post, Benjamin Van Durme |
| 2019 | NAACL | On Measuring Social Biases in Sentence Encoders. | Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, Rachel Rudinger |
| 2018 | EMNLP | Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation. | Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, Benjamin Van Durme |
| 2018 | EMNLP | Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation. | Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, Benjamin Van Durme |
| 2018 | EMNLP | Neural-Davidsonian Semantic Proto-role Labeling. | Rachel Rudinger, Adam R. Teichert, Ryan Culkin, Sheng Zhang, Benjamin Van Durme |
| 2018 | EMNLP | Lexicosyntactic inference in neural models. | Aaron Steven White, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
| 2018 | EMNLP | Cross-lingual Decompositional Semantic Parsing. | Sheng Zhang, Xutai Ma, Rachel Rudinger, Kevin Duh, Benjamin Van Durme |
| 2018 | NAACL | Gender Bias in Coreference Resolution. | Rachel Rudinger, Jason Naradowsky, Brian Leonard, Benjamin Van Durme |
| 2018 | NAACL | Neural Models of Factuality. | Rachel Rudinger, Aaron Steven White, Benjamin Van Durme |
| 2016 | EMNLP | Universal Decompositional Semantics on Universal Dependencies. | Aaron Steven White, Dee Ann Reisinger, Keisuke Sakaguchi, Tim Vieira, Sheng Zhang, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
| 2015 | EMNLP | Script Induction as Language Modeling. | Rachel Rudinger, Pushpendre Rastogi, Francis Ferraro, Benjamin Van Durme |
| 2013 | ACL | SenseSpotting: Never let your parallel data tie you to an old domain. | Marine Carpuat, Hal Daum III, Katharine Henry, Ann Irvine, Jagadeesh Jagarlamudi, Rachel Rudinger |