Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering.
Hwan Chang, Yumin Kim, Yonghyun Jun, Hwanhee Lee
Browse the full EMNLP paper archive.
Hwan Chang, Yumin Kim, Yonghyun Jun, Hwanhee Lee
Browse the full EMNLP paper archive.