Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller
Browse the full ICLR paper archive.
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller
Browse the full ICLR paper archive.