"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak.
Lingrui Mei, Shenghua Liu, Yiwei Wang, Baolong Bi, Jiayi Mao, Xueqi Cheng
Browse the full COLING paper archive.
Lingrui Mei, Shenghua Liu, Yiwei Wang, Baolong Bi, Jiayi Mao, Xueqi Cheng
Browse the full COLING paper archive.