Paul Rttger
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
29
Venues
8
Active years
2021–2026
Best venue rank
A*
Where they publish
Papers
29 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2026 | EACL | Bias in the East, Bias in the West: A Bilingual Analysis of LLM Political Bias on U.S.- and China-Related Issues. | Ying Ying Lim, Paul Rttger |
| 2026 | EACL | The Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models. | Giuseppe Russo, Debora Nozza, Paul Rttger, Dirk Hovy |
| 2026 | LREC | This House Debates AI: Evaluating a Language Model in Oxford-Style Debates against Human Experts. | Umberto Belluzzo, Kobi Hackenburg, Hannah Rose Kirk, Scott Hale, Paul Rttger |
| 2025 | AAAI | SafetyPrompts: A Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety. | Paul Rttger, Fabio Pernisi, Bertie Vidgen, Dirk Hovy |
| 2025 | ACL | Around the World in 24 Hours: Probing LLM Knowledge of Time and Place. | Carolin Holtermann, Paul Rttger, Anne Lauscher |
| 2025 | ACL | Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions. | Matthias Orlikowski, Jiaxin Pei, Paul Rttger, Philipp Cimiano, David Jurgens, Dirk Hovy |
| 2025 | ACL | HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter. | Manuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale, Samuel Fraiberger, Vctor Orozco-Olvera, Paul Rttger |
| 2025 | EMNLP | Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance. | Pedro Henrique Luz de Araujo, Paul Rttger, Dirk Hovy, Benjamin Roth |
| 2025 | EMNLP | TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent. | Dominik Meier, Jan Philip Wahle, Paul Rttger, Terry Ruas, Bela Gipp |
| 2025 | EMNLP | Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them. | Emanuele Moscato, Tiancheng Hu, Matthias Orlikowski, Paul Rttger, Debora Nozza |
| 2025 | ICLR | Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation. | Xinpeng Wang, Chengzhi Hu, Paul Rttger, Barbara Plank |
| 2025 | NAACL | Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations. | Yong Cao, Haijiang Liu, Arnav Arora, Isabelle Augenstein, Paul Rttger, Daniel Hershcovich |
| 2025 | NAACL | AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages. | Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, David Ifeoluwa Adelani, Ibrahim Said Ahmad, Saminu Mohammad Aliyu, Paul Rttger, Abigail Oppong, Andiswa Bukula, Chiamaka Ijeoma Chukwuneke, Ebrahim Chekol Jibril, Elyas Abdi Ismail, Esubalew Alemneh, Hagos Tesfahun Gebremichael, Lukman Jibril Aliyu, Meriem Beloucif, Oumaima Hourrane, Rooweither Mabuya, Salomey Osei, Samuel Rutunda, Tadesse Destaw Belay, Tadesse Kebede Guge, Tesfa Tegegne Asfaw, Lilian Diana Awuor Wanzare, Nelson Odhiambo Onyango, Seid Muhie Yimam, Nedjma Ousidhoum |
| 2024 | ACL | "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models. | Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Rttger, Frauke Kreuter, Dirk Hovy, Barbara Plank |
| 2024 | ACL | Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ. | Carolin Holtermann, Paul Rttger, Timm Dill, Anne Lauscher |
| 2024 | ACL | Compromesso! Italian Many-Shot Jailbreaks undermine the safety of Large Language Models. | Fabio Pernisi, Dirk Hovy, Paul Rttger |
| 2024 | ACL | Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models. | Paul Rttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schtze, Dirk Hovy |
| 2024 | ICLR | Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions. | Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Rttger, Dan Jurafsky, Tatsunori Hashimoto, James Zou |
| 2024 | ICML | Position: Near to Mid-term Risks and Opportunities of Open-Source Generative AI. | Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schrder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Rttger, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, Jakob N. Foerster |
| 2024 | NAACL | Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset. | Janis Goldzycher, Paul Rttger, Gerold Schneider |
| 2024 | NAACL | XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. | Paul Rttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, Dirk Hovy |
| 2023 | ACL | Improving the Detection of Multilingual Online Attacks with Rich Social Media Data from Singapore. | Janosch Haber, Bertie Vidgen, Matthew Chapman, Vibhor Agarwal, Roy Ka-Wei Lee, Yong Keong Yap, Paul Rttger |
| 2023 | ACL | The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics. | Matthias Orlikowski, Paul Rttger, Philipp Cimiano, Dirk Hovy |
| 2023 | EMNLP | The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values. | Hannah Kirk, Andrew M. Bean, Bertie Vidgen, Paul Rttger, Scott Hale |
| 2022 | EMNLP | Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages. | Paul Rttger, Debora Nozza, Federico Bianchi, Dirk Hovy |
| 2022 | NAACL | Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate. | Hannah Kirk, Bertie Vidgen, Paul Rttger, Tristan Thrush, Scott Hale |
| 2022 | NAACL | Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks. | Paul Rttger, Bertie Vidgen, Dirk Hovy, Janet B. Pierrehumbert |
| 2021 | ACL | HateCheck: Functional Tests for Hate Speech Detection Models. | Paul Rttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Z. Margetts, Janet B. Pierrehumbert |
| 2021 | EMNLP | Temporal Adaptation of BERT and Performance on Downstream Document Classification: Insights from Social Media. | Paul Rttger, Janet B. Pierrehumbert |