Reinforcement Learning-powered Effectiveness and Efficiency Few-shot Jailbreaking Attack LLMs.
Xuehai Tang, Zhongjiang Yao, Jie Wen, Yangchen Dong, Jizhong Han, Songlin Hu
Browse the full ISPA paper archive.
Xuehai Tang, Zhongjiang Yao, Jie Wen, Yangchen Dong, Jizhong Han, Songlin Hu
Browse the full ISPA paper archive.