| 2025 | ICLR | System 1.x: Learning to Balance Fast and Slow Planning with Language Models. | Swarnadeep Saha, Archiki Prasad, Justin Chih-Yao Chen, Peter Hase, Elias Stengel-Eskin, Mohit Bansal |
| 2025 | NAACL | Teaching Models to Balance Resisting and Accepting Persuasion. | Elias Stengel-Eskin, Peter Hase, Mohit Bansal |
| 2024 | ACL | The Unreasonable Effectiveness of Easy Training Data for Hard Tasks. | Peter Hase, Mohit Bansal, Peter Clark, Sarah Wiegreffe |
| 2024 | ICLR | Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks. | Vaidehi Patil, Peter Hase, Mohit Bansal |
| 2023 | EACL | Methods for Measuring, Updating, and Visualizing Factual Beliefs in Language Models. | Peter Hase, Mona T. Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer |
| 2023 | EACL | GrIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models. | Archiki Prasad, Peter Hase, Xiang Zhou, Mohit Bansal |
| 2023 | ICLR | Summarization Programs: Interpretable Abstractive Summarization with Neural Modular Trees. | Swarnadeep Saha, Shiyue Zhang, Peter Hase, Mohit Bansal |
| 2022 | EMNLP | Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated Explanations. | Swarnadeep Saha, Peter Hase, Nazneen Rajani, Mohit Bansal |
| 2021 | EMNLP | FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging. | Han Guo, Nazneen Rajani, Peter Hase, Mohit Bansal, Caiming Xiong |
| 2020 | ACL | Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior? | Peter Hase, Mohit Bansal |
| 2020 | EMNLP | Leakage-Adjusted Simulatability: Can Models Generate Non-Trivial Explanations of Their Behavior in Natural Language? | Peter Hase, Shiyue Zhang, Harry Xie, Mohit Bansal |
| 2019 | HCOMP | Interpretable Image Recognition with Hierarchical Prototypes. | Peter Hase, Chaofan Chen, Oscar Li, Cynthia Rudin |