Skip to content

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.