Skip to content

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation.

Jun Zhuang, Hai Jin, Ye Zhang, Zhengjian Kang, Wenbin Zhang, Gaby G. Dagher, Haohan Wang

VenueA*EMNLP
Year2025
ProceedingsEMNLP (Findings)

Browse the full EMNLP paper archive.