Interpretable Reward Model via Sparse Autoencoder.
Shuyi Zhang, Wei Shi, Sihang Li, Jiayi Liao, Tao Liang, Hengxing Cai, Xiang Wang
Browse the full AAAI paper archive.
Shuyi Zhang, Wei Shi, Sihang Li, Jiayi Liao, Tao Liang, Hengxing Cai, Xiang Wang
Browse the full AAAI paper archive.