Denevil: towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning.
Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, Ning Gu
Browse the full ICLR paper archive.
Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, Ning Gu
Browse the full ICLR paper archive.