A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models.
Dong Shu, Xuansheng Wu, Haiyan Zhao, Daking Rai, Ziyu Yao, Ninghao Liu, Mengnan Du
Browse the full EMNLP paper archive.
Dong Shu, Xuansheng Wu, Haiyan Zhao, Daking Rai, Ziyu Yao, Ninghao Liu, Mengnan Du
Browse the full EMNLP paper archive.