Skip to content

Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack.

Yu Fu, Yufei Li, Wen Xiao, Cong Liu, Yue Dong

VenueA*ACL
Year2024
ProceedingsACL (1)

Browse the full ACL paper archive.