SafeInt: Shielding Large Language Models from Jailbreak Attacks via Safety-Aware Representation Intervention.
Jiaqi Wu, Chen Chen, Chunyan Hou, Xiaojie Yuan
Browse the full EMNLP paper archive.
Jiaqi Wu, Chen Chen, Chunyan Hou, Xiaojie Yuan
Browse the full EMNLP paper archive.