| 2025 | ICLR | VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models. | Lisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez |
| 2025 | ICLR | Monitoring Latent World States in Language Models with Propositional Probes. | Jiahai Feng, Stuart Russell, Jacob Steinhardt |
| 2025 | ICLR | Interpreting the Second-Order Effects of Neurons in CLIP. | Yossi Gandelsman, Alexei A. Efros, Jacob Steinhardt |
| 2025 | ICLR | Uncovering Gaps in How Humans and LLMs Interpret Subjective Language. | Erik Jones, Arjun Patrawala, Jacob Steinhardt |
| 2025 | ICLR | Language Models Learn to Mislead Humans via RLHF. | Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He, Shi Feng |
| 2025 | ICLR | Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision. | Yaowen Ye, Cassidy Laidlaw, Jacob Steinhardt |
| 2025 | ICML | Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts. | Jiahai Feng, Stuart Russell, Jacob Steinhardt |
| 2025 | ICML | Adversaries Can Misuse Combinations of Safe Models. | Erik Jones, Anca D. Dragan, Jacob Steinhardt |
| 2025 | ICML | What Do Learning Dynamics Reveal About Generalization in LLM Mathematical Reasoning? | Katie Kang, Amrith Setlur, Dibya Ghosh, Jacob Steinhardt, Claire J. Tomlin, Sergey Levine, Aviral Kumar |
| 2025 | ICML | Eliciting Language Model Behaviors with Investigator Agents. | Xiang Lisa Li, Neil Chowdhury, Daniel D. Johnson, Tatsunori Hashimoto, Percy Liang, Sarah Schwettmann, Jacob Steinhardt |
| 2025 | ICML | Which Attention Heads Matter for In-Context Learning? | Kayo Yin, Jacob Steinhardt |
| 2024 | CVPR | Describing Differences in Image Sets with Natural Language. | Lisa Dunlap, Yuhui Zhang, Xiaohan Wang, Ruiqi Zhong, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez, Serena Yeung-Levy |
| 2024 | ICLR | How do Language Models Bind Entities in Context? | Jiahai Feng, Jacob Steinhardt |
| 2024 | ICLR | Interpreting CLIP's Image Representation via Text-Based Decomposition. | Yossi Gandelsman, Alexei A. Efros, Jacob Steinhardt |
| 2024 | ICLR | Overthinking the Truth: Understanding how Language Models Process False Demonstrations. | Danny Halawi, Jean-Stanislas Denain, Jacob Steinhardt |
| 2024 | ICML | Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations. | Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao, He He, Jacob Steinhardt, Zhou Yu, Kathleen R. McKeown |
| 2024 | ICML | Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation. | Danny Halawi, Alexander Wei, Eric Wallace, Tony Tong Wang, Nika Haghtalab, Jacob Steinhardt |
| 2024 | ICML | Feedback Loops With Language Models Drive In-Context Reward Hacking. | Alexander Pan, Erik Jones, Meena Jagadeesan, Jacob Steinhardt |
| 2023 | AISTATS | Reward Learning as Doubly Nonparametric Bandits: Optimal Design and Scaling Laws. | Kush Bhatia, Wenshuo Guo, Jacob Steinhardt |
| 2023 | ICLR | Discovering Latent Knowledge in Language Models Without Supervision. | Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt |
| 2023 | ICLR | Progress measures for grokking via mechanistic interpretability. | Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, Jacob Steinhardt |
| 2023 | ICLR | Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small. | Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, Jacob Steinhardt |
| 2023 | ICML | Automatically Auditing Large Language Models via Discrete Optimization. | Erik Jones, Anca D. Dragan, Aditi Raghunathan, Jacob Steinhardt |
| 2023 | ICML | Are Neurons Actually Collapsed? On the Fine-Grained Structure in Neural Representations. | Yongyi Yang, Jacob Steinhardt, Wei Hu |
| 2022 | CVPR | PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures. | Dan Hendrycks, Andy Zou, Mantas Mazeika, Leonard Tang, Bo Li, Dawn Song, Jacob Steinhardt |
| 2022 | CVPR | A3D: Studying Pretrained Representations with Programmable Datasets. | Ye Wang, Norman Mu, Daniele Grandi, Nicolas Savva, Jacob Steinhardt |
| 2022 | ICLR | The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. | Alexander Pan, Kush Bhatia, Jacob Steinhardt |
| 2022 | ICML | More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations Generalize. | Alexander Wei, Wei Hu, Jacob Steinhardt |
| 2022 | ICML | Scaling Out-of-Distribution Detection for Real-World Settings. | Dan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou, Joseph Kwon, Mohammadreza Mostajabi, Jacob Steinhardt, Dawn Song |
| 2022 | ICML | Predicting Out-of-Distribution Error with the Projection Norm. | Yaodong Yu, Zitong Yang, Alexander Wei, Yi Ma, Jacob Steinhardt |
| 2022 | ICML | Describing Differences between Text Distributions with Natural Language. | Ruiqi Zhong, Charlie Snell, Dan Klein, Jacob Steinhardt |
| 2021 | ACL | Are Larger Pretrained Language Models Uniformly Better? Comparing Performance at the Instance Level. | Ruiqi Zhong, Dhruba Ghosh, Dan Klein, Jacob Steinhardt |
| 2021 | CVPR | Limitations of Post-Hoc Feature Alignment for Robustness. | Collin Burns, Jacob Steinhardt |
| 2021 | CVPR | Natural Adversarial Examples. | Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, Dawn Song |
| 2021 | ICCV | The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization. | Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, Justin Gilmer |
| 2021 | ICLR | Aligning AI With Shared Human Values. | Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, Jacob Steinhardt |
| 2021 | ICLR | Measuring Massive Multitask Language Understanding. | Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, Jacob Steinhardt |
| 2020 | ICML | Identifying Statistical Bias in Dataset Replication. | Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Jacob Steinhardt, Aleksander Madry |
| 2020 | ICML | Rethinking Bias-Variance Trade-off for Generalization of Neural Networks. | Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, Yi Ma |
| 2020 | ISIT | When does the Tukey Median work? | Banghua Zhu, Jiantao Jiao, Jacob Steinhardt |
| 2019 | ICML | Sever: A Robust Meta-Algorithm for Stochastic Optimization. | Ilias Diakonikolas, Gautam Kamath, Daniel Kane, Jerry Li, Jacob Steinhardt, Alistair Stewart |
| 2018 | ICLR | Certified Defenses against Adversarial Examples. | Aditi Raghunathan, Jacob Steinhardt, Percy Liang |
| 2018 | STOC | Robust moment estimation and improved clustering via sum of squares. | Pravesh K. Kothari, Jacob Steinhardt, David Steurer |
| 2017 | STOC | Learning from untrusted data. | Moses Charikar, Jacob Steinhardt, Gregory Valiant |
| 2016 | COLT | Memory, Communication, and Statistical Queries. | Jacob Steinhardt, Gregory Valiant, Stefan Wager |
| 2015 | AISTATS | Learning Where to Sample in Structured Prediction. | Tianlin Shi, Jacob Steinhardt, Percy Liang |
| 2015 | COLT | Minimax rates for memory-bounded sparse linear regression. | Jacob Steinhardt, John C. Duchi |
| 2015 | ICML | Reified Context Models. | Jacob Steinhardt, Percy Liang |
| 2015 | ICML | Learning Fast-Mixing Models for Structured Prediction. | Jacob Steinhardt, Percy Liang |
| 2014 | ICML | Filtering with Abstract Particles. | Jacob Steinhardt, Percy Liang |
| 2014 | ICML | Adaptivity and Optimism: An Improved Exponentiated Gradient Algorithm. | Jacob Steinhardt, Percy Liang |