Feint and Attack: Jailbreaking and Protecting LLMs via Attention Distribution Modeling.
Rui Pu, Chaozhuo Li, Rui Ha, Zejian Chen, Litian Zhang, Zheng Liu, Lirong Qiu, Zaisheng Ye
Browse the full IJCAI paper archive.
Rui Pu, Chaozhuo Li, Rui Ha, Zejian Chen, Litian Zhang, Zheng Liu, Lirong Qiu, Zaisheng Ye
Browse the full IJCAI paper archive.