| 2025 | ICLR | Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models. | Javier Ferrando, Oscar Balcells Obeso, Senthooran Rajamanoharan, Neel Nanda |
| 2025 | ICLR | Sparse Autoencoders Do Not Find Canonical Units of Analysis. | Patrick Leask, Bart Bussmann, Michael T. Pearce, Joseph Isaac Bloom, Curt Tigges, Noura Al Moubayed, Lee Sharkey, Neel Nanda |
| 2025 | ICLR | Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control. | Aleksandar Makelov, Georg Lange, Neel Nanda |
| 2025 | ICML | Learning Multi-Level Features with Matryoshka Sparse Autoencoders. | Bart Bussmann, Noa Nabeshima, Adam Karvonen, Neel Nanda |
| 2025 | ICML | Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. | Subhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark, Neel Nanda |
| 2025 | ICML | SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability. | Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Isaac Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Demian Till, Matthew Wearden, Arthur Conmy, Samuel Marks, Neel Nanda |
| 2025 | ICML | Scaling Sparse Feature Circuits For Studying In-Context Learning. | Dmitrii Kharlapenko, Stepan Shabalin, Arthur Conmy, Neel Nanda |
| 2025 | ICML | Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models. | Patrick Leask, Neel Nanda, Noura Al Moubayed |
| 2024 | ICLR | Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching. | Aleksandar Makelov, Georg Lange, Atticus Geiger, Neel Nanda |
| 2024 | ICLR | Towards Best Practices of Activation Patching in Language Models: Metrics and Methods. | Fred Zhang, Neel Nanda |
| 2024 | ICML | Explorations of Self-Repair in Language Models. | Cody Rushing, Neel Nanda |
| 2023 | ICLR | Progress measures for grokking via mechanistic interpretability. | Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, Jacob Steinhardt |
| 2023 | ICML | A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations. | Bilal Chughtai, Lawrence Chan, Neel Nanda |