How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions.
Lorenzo Pacchiardi, Alex James Chan, Sren Mindermann, Ilan Moscovitz, Alexa Y. Pan, Yarin Gal, Owain Evans, Jan Markus Brauner
Browse the full ICLR paper archive.