Skip to content

Owain Evans

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

7

Venues

4

Active years

2016–2025

Best venue rank

A*

Where they publish

Papers

7 indexed papers, newest first.

YearVenueTitleAuthors
2025ICLRTell me about yourself: LLMs are aware of their learned behaviors.Jan Betley, Xuchan Bao, Martn Soto, Anna Sztyber-Betley, James Chua, Owain Evans
2025ICLRLooking Inward: Language Models Can Learn About Themselves by Introspection.Felix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, Owain Evans
2025ICMLEmergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs.Jan Betley, Daniel Chee Hian Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martn Soto, Nathan Labenz, Owain Evans
2024ICLRThe Reversal Curse: LLMs trained on "A is B" fail to learn "B is A".Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, Owain Evans
2024ICLRHow to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions.Lorenzo Pacchiardi, Alex James Chan, Sren Mindermann, Ilan Moscovitz, Alexa Y. Pan, Yarin Gal, Owain Evans, Jan Markus Brauner
2022ACLTruthfulQA: Measuring How Models Mimic Human Falsehoods.Stephanie Lin, Jacob Hilton, Owain Evans
2016AAAILearning the Preferences of Ignorant, Inconsistent Agents.Owain Evans, Andreas Stuhlmller, Noah D. Goodman