Skip to content

Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, Lee Sharkey

VenueA*ICLR
Year2024
ProceedingsICLR

Browse the full ICLR paper archive.