Skip to content

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations.

Xiaohu Li, Yunfeng Ning, Zepeng Bao, Mayi Xu, Jianhao Chen, Tieyun Qian

VenueA*ACL
Year2025
ProceedingsACL (Findings)

Browse the full ACL paper archive.