Skip to content

Atticus Geiger

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

20

Venues

6

Active years

2019–2026

Best venue rank

A*

Where they publish

Papers

20 indexed papers, newest first.

YearVenueTitleAuthors
2026ACLConstructing Interpretable Features from Compositional Neuron Groups.Or David Shafran, Atticus Geiger, Mor Geva
2025ACLEnhancing Automated Interpretability with Output-Centric Feature Descriptions.Yoav Gur-Arieh, Roy Mayan, Chen Agassy, Atticus Geiger, Mor Geva
2025ICLRHyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks.Jiuding Sun, Jing Huang, Sidharth Baskaran, Karel D'Oosterlinck, Christopher Potts, Michael Sklar, Atticus Geiger
2025ICMLMIB: A Mechanistic Interpretability Benchmark.Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Ivn Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fried Fiotto-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov
2025ICMLAxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders.Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D. Manning, Christopher Potts
2025ICMLHow Do Transformers Learn Variable Binding in Symbolic Programs?Yiwei Wu, Atticus Geiger, Raphal Millire
2024ACLRAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations.Jing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva, Atticus Geiger
2024EMNLPUpdating CLIP to Prefer Descriptions Over Captions.Amir Zur, Elisa Kreiss, Karel D'Oosterlinck, Christopher Potts, Atticus Geiger
2024ICLRIs This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching.Aleksandar Makelov, Georg Lange, Atticus Geiger, Neel Nanda
2024NAACLpyvene: A Library for Understanding and Improving PyTorch Models via Interventions.Zhengxuan Wu, Atticus Geiger, Aryaman Arora, Jing Huang, Zheng Wang, Noah D. Goodman, Christopher D. Manning, Christopher Potts
2023ACLScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning.Jingyuan Selena She, Christopher Potts, Samuel R. Bowman, Atticus Geiger
2023CogSciA Semantics for Causing, Enabling, and Preventing Verbs Using Structural Causal Models.Angela Cao, Atticus Geiger, Elisa Kreiss, Thomas Icard, Tobias Gerstenberg
2023ICMLCausal Proxy Models for Concept-based Model Explanations.Zhengxuan Wu, Karel D'Oosterlinck, Atticus Geiger, Amir Zur, Christopher Potts
2022ICMLInducing Causal Structure for Interpretable Neural Networks.Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah D. Goodman, Christopher Potts
2022NAACLCausal Distillation for Language Models.Zhengxuan Wu, Atticus Geiger, Joshua Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, Noah D. Goodman
2021ACLDynaSent: A Dynamic Benchmark for Sentiment Analysis.Christopher Potts, Zhengxuan Wu, Atticus Geiger, Douwe Kiela
2021NAACLDynabench: Rethinking Benchmarking in NLP.Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams
2020CogSciRelational reasoning and generalization using non-symbolic neural networks.Atticus Geiger, Alexandra Carstensen, Michael C. Frank, Christopher Potts
2019EMNLPPosing Fair Generalization Tasks for Natural Language Inference.Atticus Geiger, Ignacio Cases, Lauri Karttunen, Christopher Potts
2019NAACLRecursive Routing Networks: Learning to Compose Modules for Language Understanding.Ignacio Cases, Clemens Rosenbaum, Matthew Riemer, Atticus Geiger, Tim Klinger, Alex Tamkin, Olivia Li, Sandhini Agarwal, Joshua D. Greene, Dan Jurafsky, Christopher Potts, Lauri Karttunen