The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models.
Yan Liu, Yu Liu, Xiaokang Chen, Pin-Yu Chen, Daoguang Zan, Min-Yen Kan, Tsung-Yi Ho
Browse the full ICLR paper archive.
Yan Liu, Yu Liu, Xiaokang Chen, Pin-Yu Chen, Daoguang Zan, Min-Yen Kan, Tsung-Yi Ho
Browse the full ICLR paper archive.