| 2024 | ACL | Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models. | Paul Rttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schtze, Dirk Hovy |
| 2024 | NAACL | XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. | Paul Rttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, Dirk Hovy |
| 2023 | EMNLP | The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values. | Hannah Kirk, Andrew M. Bean, Bertie Vidgen, Paul Rttger, Scott Hale |
| 2022 | EMNLP | Handling and Presenting Harmful Text in NLP Research. | Hannah Kirk, Abeba Birhane, Bertie Vidgen, Leon Derczynski |
| 2022 | IJCNLP | A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning. | Hugo Berg, Siobhan Mackenzie Hall, Yash Bhalgat, Hannah Kirk, Aleksandar Shtedritski, Max Bain |
| 2022 | NAACL | Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate. | Hannah Kirk, Bertie Vidgen, Paul Rttger, Tristan Thrush, Scott Hale |