Skip to content

Aaron Mueller

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

34

Venues

9

Active years

2019–2026

Best venue rank

A*

Where they publish

Papers

34 indexed papers, newest first.

YearVenueTitleAuthors
2026ACLCRISP: Persistent Concept Unlearning via Sparse Autoencoders.Tomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek, Yonatan Belinkov
2026ACLCrosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining.Deniz Bayazit, Aaron Mueller, Antoine Bosselut
2026ACLFrom Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?Aaron Mueller, Andrew Lee, Shruti Joshi, Ekdeep Singh Lubana, Dhanya Sridhar, Patrik Reizinger
2026ACLLatent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate.John Seon Keun Yi, Aaron Mueller, Dokyun Lee
2026EACLMeasuring Mechanistic Independence: Can Bias Be Removed Without Erasing Demographics?Zhengyang Shan, Aaron Mueller
2026EACLImproving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection.Joe Stacey, Lisa Alazraki, Aran Ubhi, Beyza Ermis, Aaron Mueller, Marek Rei
2025ACLPosition-aware Automatic Circuit Discovery.Tal Haklay, Hadas Orgad, David Bau, Aaron Mueller, Yonatan Belinkov
2025EMNLPSAEs Are Good for Steering - If You Select the Right Features.Dana Arad, Aaron Mueller, Yonatan Belinkov
2025ICLRNNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals.Jaden Fried Fiotto-Kaufman, Alexander Russell Loftus, Eric Todd, Jannik Brinkmann, Koyena Pal, Dmitrii Troitskii, Michael Ripa, Adam Belfki, Can Rager, Caden Juang, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Nikhil Prakash, Carla E. Brodley, Arjun Guha, Jonathan Bell, Byron C. Wallace, David Bau
2025ICLRSparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller
2025ICLRArithmetic Without Algorithms: Language Models Solve Math with a Bag of Heuristics.Yaniv Nikankin, Anja Reusch, Aaron Mueller, Yonatan Belinkov
2025ICMLMIB: A Mechanistic Interpretability Benchmark.Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Ivn Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fried Fiotto-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov
2025NAACLLarge Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages.Jannik Brinkmann, Chris Wendler, Christian Bartelt, Aaron Mueller
2025NAACLIncremental Sentence Processing Mechanisms in Autoregressive Transformer Language Models.Michael Hanna, Aaron Mueller
2025NAACLCharacterizing the Role of Similarity in the Property Inferences of Language Models.Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra
2024CogSciInsights from the first BabyLM Challenge: Training sample-efficient language models on a developmentally plausible corpus.Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Gotlieb Wilcox, Chengxu Zhuang, Adina Williams, Ryan Cotterell, Tal Linzen
2024ICLRFunction Vectors in Large Language Models.Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, David Bau
2024NAACLIn-context Learning Generalizes, But Not Always Robustly: The Case of Syntax.Aaron Mueller, Albert Webson, Jackson Petty, Tal Linzen
2023ACLWhat Do NLP Researchers Believe? Results of the NLP Community Metasurvey.Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman
2023ACLHow to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive Biases.Aaron Mueller, Tal Linzen
2023ACLMeta-training with Demonstration Retrieval for Efficient Few-shot Learning.Aaron Mueller, Kanika Narang, Lambert Mathias, Qifan Wang, Hamed Firooz
2023ACLLanguage model acceptability judgements are not always robust to context.Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams
2022ACLColoring the Blank Slate: Pre-training Imparts a Hierarchical Inductive Bias to Sequence-to-sequence Models.Aaron Mueller, Robert Frank, Tal Linzen, Luheng Wang, Sebastian Schuster
2022ACLLabel Semantic Aware Pre-training for Few-shot Text Classification.Aaron Mueller, Jason Krone, Salvatore Romeo, Saab Mansour, Elman Mansimov, Yi Zhang, Dan Roth
2022CoNLLCausal Analysis of Syntactic Agreement Neurons in Multilingual Language Models.Aaron Mueller, Yu Xia, Tal Linzen
2022EMNLPBernice: A Multilingual Pre-trained Encoder for Twitter.Alexandra DeLucia, Shijie Wu, Aaron Mueller, Carlos Alejandro Aguirre, Philip Resnik, Mark Dredze
2021ACLCausal Analysis of Syntactic Agreement Mechanisms in Neural Language Models.Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber, Tal Linzen, Yonatan Belinkov
2021NAACLFine-tuning Encoders for Improved Monolingual and Zero-shot Polylingual Neural Topic Modeling.Aaron Mueller, Mark Dredze
2020ACLCross-Linguistic Syntactic Evaluation of Word Prediction Models.Aaron Mueller, Garrett Nicolai, Panayiota Petrou-Zeniou, Natalia Talmina, Tal Linzen
2020LRECThe Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration.Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, David Yarowsky
2020LRECAn Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages.Aaron Mueller, Garrett Nicolai, Arya D. McCarthy, Dylan Lewis, Winston Wu, David Yarowsky
2020LRECFine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages.Garrett Nicolai, Dylan Lewis, Arya D. McCarthy, Aaron Mueller, Winston Wu, David Yarowsky
2019EMNLPModeling Color Terminology Across Thousands of Languages.Arya D. McCarthy, Winston Wu, Aaron Mueller, Bill Watson, David Yarowsky
2019EMNLPQuantity doesn't buy quality syntax with neural language models.Marten van Schijndel, Aaron Mueller, Tal Linzen