Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak Attacks.
Yue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang Zhang
Browse the full EMNLP paper archive.
Yue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang Zhang
Browse the full EMNLP paper archive.