Skip to content

BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models.

Yi Zeng, Weiyu Sun, Tran Ngoc Huynh, Dawn Song, Bo Li, Ruoxi Jia

VenueA*EMNLP
Year2024
ProceedingsEMNLP

Browse the full EMNLP paper archive.