Skip to content

Neel Nanda

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

13

Venues

2

Active years

2023–2025

Best venue rank

A*

Where they publish

Papers

13 indexed papers, newest first.

YearVenueTitleAuthors
2025ICLRDo I Know This Entity? Knowledge Awareness and Hallucinations in Language Models.Javier Ferrando, Oscar Balcells Obeso, Senthooran Rajamanoharan, Neel Nanda
2025ICLRSparse Autoencoders Do Not Find Canonical Units of Analysis.Patrick Leask, Bart Bussmann, Michael T. Pearce, Joseph Isaac Bloom, Curt Tigges, Noura Al Moubayed, Lee Sharkey, Neel Nanda
2025ICLRTowards Principled Evaluations of Sparse Autoencoders for Interpretability and Control.Aleksandar Makelov, Georg Lange, Neel Nanda
2025ICMLLearning Multi-Level Features with Matryoshka Sparse Autoencoders.Bart Bussmann, Noa Nabeshima, Adam Karvonen, Neel Nanda
2025ICMLAre Sparse Autoencoders Useful? A Case Study in Sparse Probing.Subhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark, Neel Nanda
2025ICMLSAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability.Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Isaac Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Demian Till, Matthew Wearden, Arthur Conmy, Samuel Marks, Neel Nanda
2025ICMLScaling Sparse Feature Circuits For Studying In-Context Learning.Dmitrii Kharlapenko, Stepan Shabalin, Arthur Conmy, Neel Nanda
2025ICMLInference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models.Patrick Leask, Neel Nanda, Noura Al Moubayed
2024ICLRIs This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching.Aleksandar Makelov, Georg Lange, Atticus Geiger, Neel Nanda
2024ICLRTowards Best Practices of Activation Patching in Language Models: Metrics and Methods.Fred Zhang, Neel Nanda
2024ICMLExplorations of Self-Repair in Language Models.Cody Rushing, Neel Nanda
2023ICLRProgress measures for grokking via mechanistic interpretability.Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, Jacob Steinhardt
2023ICMLA Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations.Bilal Chughtai, Lawrence Chan, Neel Nanda