ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding.
Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, Dacheng Tao
Browse the full ACL paper archive.
Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, Dacheng Tao
Browse the full ACL paper archive.