Skip to content

Paul Rttger

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

29

Venues

8

Active years

2021–2026

Best venue rank

A*

Where they publish

Papers

29 indexed papers, newest first.

YearVenueTitleAuthors
2026EACLBias in the East, Bias in the West: A Bilingual Analysis of LLM Political Bias on U.S.- and China-Related Issues.Ying Ying Lim, Paul Rttger
2026EACLThe Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models.Giuseppe Russo, Debora Nozza, Paul Rttger, Dirk Hovy
2026LRECThis House Debates AI: Evaluating a Language Model in Oxford-Style Debates against Human Experts.Umberto Belluzzo, Kobi Hackenburg, Hannah Rose Kirk, Scott Hale, Paul Rttger
2025AAAISafetyPrompts: A Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety.Paul Rttger, Fabio Pernisi, Bertie Vidgen, Dirk Hovy
2025ACLAround the World in 24 Hours: Probing LLM Knowledge of Time and Place.Carolin Holtermann, Paul Rttger, Anne Lauscher
2025ACLBeyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions.Matthias Orlikowski, Jiaxin Pei, Paul Rttger, Philipp Cimiano, David Jurgens, Dirk Hovy
2025ACLHateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter.Manuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale, Samuel Fraiberger, Vctor Orozco-Olvera, Paul Rttger
2025EMNLPPrincipled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance.Pedro Henrique Luz de Araujo, Paul Rttger, Dirk Hovy, Benjamin Roth
2025EMNLPTrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent.Dominik Meier, Jan Philip Wahle, Paul Rttger, Terry Ruas, Bela Gipp
2025EMNLPPersonalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them.Emanuele Moscato, Tiancheng Hu, Matthias Orlikowski, Paul Rttger, Debora Nozza
2025ICLRSurgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation.Xinpeng Wang, Chengzhi Hu, Paul Rttger, Barbara Plank
2025NAACLSpecializing Large Language Models to Simulate Survey Response Distributions for Global Populations.Yong Cao, Haijiang Liu, Arnav Arora, Isabelle Augenstein, Paul Rttger, Daniel Hershcovich
2025NAACLAfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages.Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, David Ifeoluwa Adelani, Ibrahim Said Ahmad, Saminu Mohammad Aliyu, Paul Rttger, Abigail Oppong, Andiswa Bukula, Chiamaka Ijeoma Chukwuneke, Ebrahim Chekol Jibril, Elyas Abdi Ismail, Esubalew Alemneh, Hagos Tesfahun Gebremichael, Lukman Jibril Aliyu, Meriem Beloucif, Oumaima Hourrane, Rooweither Mabuya, Salomey Osei, Samuel Rutunda, Tadesse Destaw Belay, Tadesse Kebede Guge, Tesfa Tegegne Asfaw, Lilian Diana Awuor Wanzare, Nelson Odhiambo Onyango, Seid Muhie Yimam, Nedjma Ousidhoum
2024ACL"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models.Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Rttger, Frauke Kreuter, Dirk Hovy, Barbara Plank
2024ACLEvaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ.Carolin Holtermann, Paul Rttger, Timm Dill, Anne Lauscher
2024ACLCompromesso! Italian Many-Shot Jailbreaks undermine the safety of Large Language Models.Fabio Pernisi, Dirk Hovy, Paul Rttger
2024ACLPolitical Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models.Paul Rttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schtze, Dirk Hovy
2024ICLRSafety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Rttger, Dan Jurafsky, Tatsunori Hashimoto, James Zou
2024ICMLPosition: Near to Mid-term Risks and Opportunities of Open-Source Generative AI.Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schrder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Rttger, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, Jakob N. Foerster
2024NAACLImproving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset.Janis Goldzycher, Paul Rttger, Gerold Schneider
2024NAACLXSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.Paul Rttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, Dirk Hovy
2023ACLImproving the Detection of Multilingual Online Attacks with Rich Social Media Data from Singapore.Janosch Haber, Bertie Vidgen, Matthew Chapman, Vibhor Agarwal, Roy Ka-Wei Lee, Yong Keong Yap, Paul Rttger
2023ACLThe Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics.Matthias Orlikowski, Paul Rttger, Philipp Cimiano, Dirk Hovy
2023EMNLPThe Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values.Hannah Kirk, Andrew M. Bean, Bertie Vidgen, Paul Rttger, Scott Hale
2022EMNLPData-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages.Paul Rttger, Debora Nozza, Federico Bianchi, Dirk Hovy
2022NAACLHatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate.Hannah Kirk, Bertie Vidgen, Paul Rttger, Tristan Thrush, Scott Hale
2022NAACLTwo Contrasting Data Annotation Paradigms for Subjective NLP Tasks.Paul Rttger, Bertie Vidgen, Dirk Hovy, Janet B. Pierrehumbert
2021ACLHateCheck: Functional Tests for Hate Speech Detection Models.Paul Rttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Z. Margetts, Janet B. Pierrehumbert
2021EMNLPTemporal Adaptation of BERT and Performance on Downstream Document Classification: Insights from Social Media.Paul Rttger, Janet B. Pierrehumbert