Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions.
Xiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi, Kaidi Xu, Di Huang, Xing Hu
Browse the full AAAI paper archive.
Xiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi, Kaidi Xu, Di Huang, Xing Hu
Browse the full AAAI paper archive.