Skip to content

Jacob Steinhardt

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

51

Venues

9

Active years

2014–2025

Best venue rank

A*

Where they publish

Papers

51 indexed papers, newest first.

YearVenueTitleAuthors
2025ICLRVibeCheck: Discover and Quantify Qualitative Differences in Large Language Models.Lisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez
2025ICLRMonitoring Latent World States in Language Models with Propositional Probes.Jiahai Feng, Stuart Russell, Jacob Steinhardt
2025ICLRInterpreting the Second-Order Effects of Neurons in CLIP.Yossi Gandelsman, Alexei A. Efros, Jacob Steinhardt
2025ICLRUncovering Gaps in How Humans and LLMs Interpret Subjective Language.Erik Jones, Arjun Patrawala, Jacob Steinhardt
2025ICLRLanguage Models Learn to Mislead Humans via RLHF.Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He, Shi Feng
2025ICLRIterative Label Refinement Matters More than Preference Optimization under Weak Supervision.Yaowen Ye, Cassidy Laidlaw, Jacob Steinhardt
2025ICMLExtractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts.Jiahai Feng, Stuart Russell, Jacob Steinhardt
2025ICMLAdversaries Can Misuse Combinations of Safe Models.Erik Jones, Anca D. Dragan, Jacob Steinhardt
2025ICMLWhat Do Learning Dynamics Reveal About Generalization in LLM Mathematical Reasoning?Katie Kang, Amrith Setlur, Dibya Ghosh, Jacob Steinhardt, Claire J. Tomlin, Sergey Levine, Aviral Kumar
2025ICMLEliciting Language Model Behaviors with Investigator Agents.Xiang Lisa Li, Neil Chowdhury, Daniel D. Johnson, Tatsunori Hashimoto, Percy Liang, Sarah Schwettmann, Jacob Steinhardt
2025ICMLWhich Attention Heads Matter for In-Context Learning?Kayo Yin, Jacob Steinhardt
2024CVPRDescribing Differences in Image Sets with Natural Language.Lisa Dunlap, Yuhui Zhang, Xiaohan Wang, Ruiqi Zhong, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez, Serena Yeung-Levy
2024ICLRHow do Language Models Bind Entities in Context?Jiahai Feng, Jacob Steinhardt
2024ICLRInterpreting CLIP's Image Representation via Text-Based Decomposition.Yossi Gandelsman, Alexei A. Efros, Jacob Steinhardt
2024ICLROverthinking the Truth: Understanding how Language Models Process False Demonstrations.Danny Halawi, Jean-Stanislas Denain, Jacob Steinhardt
2024ICMLDo Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations.Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao, He He, Jacob Steinhardt, Zhou Yu, Kathleen R. McKeown
2024ICMLCovert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation.Danny Halawi, Alexander Wei, Eric Wallace, Tony Tong Wang, Nika Haghtalab, Jacob Steinhardt
2024ICMLFeedback Loops With Language Models Drive In-Context Reward Hacking.Alexander Pan, Erik Jones, Meena Jagadeesan, Jacob Steinhardt
2023AISTATSReward Learning as Doubly Nonparametric Bandits: Optimal Design and Scaling Laws.Kush Bhatia, Wenshuo Guo, Jacob Steinhardt
2023ICLRDiscovering Latent Knowledge in Language Models Without Supervision.Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt
2023ICLRProgress measures for grokking via mechanistic interpretability.Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, Jacob Steinhardt
2023ICLRInterpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small.Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, Jacob Steinhardt
2023ICMLAutomatically Auditing Large Language Models via Discrete Optimization.Erik Jones, Anca D. Dragan, Aditi Raghunathan, Jacob Steinhardt
2023ICMLAre Neurons Actually Collapsed? On the Fine-Grained Structure in Neural Representations.Yongyi Yang, Jacob Steinhardt, Wei Hu
2022CVPRPixMix: Dreamlike Pictures Comprehensively Improve Safety Measures.Dan Hendrycks, Andy Zou, Mantas Mazeika, Leonard Tang, Bo Li, Dawn Song, Jacob Steinhardt
2022CVPRA3D: Studying Pretrained Representations with Programmable Datasets.Ye Wang, Norman Mu, Daniele Grandi, Nicolas Savva, Jacob Steinhardt
2022ICLRThe Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.Alexander Pan, Kush Bhatia, Jacob Steinhardt
2022ICMLMore Than a Toy: Random Matrix Models Predict How Real-World Neural Representations Generalize.Alexander Wei, Wei Hu, Jacob Steinhardt
2022ICMLScaling Out-of-Distribution Detection for Real-World Settings.Dan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou, Joseph Kwon, Mohammadreza Mostajabi, Jacob Steinhardt, Dawn Song
2022ICMLPredicting Out-of-Distribution Error with the Projection Norm.Yaodong Yu, Zitong Yang, Alexander Wei, Yi Ma, Jacob Steinhardt
2022ICMLDescribing Differences between Text Distributions with Natural Language.Ruiqi Zhong, Charlie Snell, Dan Klein, Jacob Steinhardt
2021ACLAre Larger Pretrained Language Models Uniformly Better? Comparing Performance at the Instance Level.Ruiqi Zhong, Dhruba Ghosh, Dan Klein, Jacob Steinhardt
2021CVPRLimitations of Post-Hoc Feature Alignment for Robustness.Collin Burns, Jacob Steinhardt
2021CVPRNatural Adversarial Examples.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, Dawn Song
2021ICCVThe Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, Justin Gilmer
2021ICLRAligning AI With Shared Human Values.Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, Jacob Steinhardt
2021ICLRMeasuring Massive Multitask Language Understanding.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, Jacob Steinhardt
2020ICMLIdentifying Statistical Bias in Dataset Replication.Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Jacob Steinhardt, Aleksander Madry
2020ICMLRethinking Bias-Variance Trade-off for Generalization of Neural Networks.Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, Yi Ma
2020ISITWhen does the Tukey Median work?Banghua Zhu, Jiantao Jiao, Jacob Steinhardt
2019ICMLSever: A Robust Meta-Algorithm for Stochastic Optimization.Ilias Diakonikolas, Gautam Kamath, Daniel Kane, Jerry Li, Jacob Steinhardt, Alistair Stewart
2018ICLRCertified Defenses against Adversarial Examples.Aditi Raghunathan, Jacob Steinhardt, Percy Liang
2018STOCRobust moment estimation and improved clustering via sum of squares.Pravesh K. Kothari, Jacob Steinhardt, David Steurer
2017STOCLearning from untrusted data.Moses Charikar, Jacob Steinhardt, Gregory Valiant
2016COLTMemory, Communication, and Statistical Queries.Jacob Steinhardt, Gregory Valiant, Stefan Wager
2015AISTATSLearning Where to Sample in Structured Prediction.Tianlin Shi, Jacob Steinhardt, Percy Liang
2015COLTMinimax rates for memory-bounded sparse linear regression.Jacob Steinhardt, John C. Duchi
2015ICMLReified Context Models.Jacob Steinhardt, Percy Liang
2015ICMLLearning Fast-Mixing Models for Structured Prediction.Jacob Steinhardt, Percy Liang
2014ICMLFiltering with Abstract Particles.Jacob Steinhardt, Percy Liang
2014ICMLAdaptivity and Optimism: An Improved Exponentiated Gradient Algorithm.Jacob Steinhardt, Percy Liang