When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations.
Huaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang, Ruixiang Tang
Browse the full ACL paper archive.
Huaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang, Ruixiang Tang
Browse the full ACL paper archive.