Skip to content

Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!

Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu, Chao Yang, Wanli Ouyang, Yu Qiao

VenueA*ACL
Year2024
ProceedingsACL (1)

Browse the full ACL paper archive.