Skip to content

Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness.

Qi Zhang, Yifei Wang, Jingyi Cui, Xiang Pan, Qi Lei, Stefanie Jegelka, Yisen Wang

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.