| 2026 | ACL | Responsible Evaluation of AI for Mental Health. | Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza del Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych |
| 2026 | EACL | Can Reasoning Help Large Language Models Capture Human Annotator Disagreement? | Jingwei Ni, Yu Fan, Vilm Zouhar, Donya Rooein, Alexander Miserlis Hoyle, Mrinmaya Sachan, Markus Leippold, Dirk Hovy, Elliott Ash |
| 2026 | EACL | PATS: Personality-Aware Teaching Strategies with Large Language Model Tutors. | Donya Rooein, Sankalan Pal Chowdhury, Mariia Eremeeva, Yuan Qin, Debora Nozza, Mrinmaya Sachan, Dirk Hovy |
| 2026 | EACL | The Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models. | Giuseppe Russo, Debora Nozza, Paul Rttger, Dirk Hovy |
| 2026 | LREC | ACID: On the Perception of Online Classism. | Arianna Muti, Elisa Bassignana, Amanda Cercas Curry, Federica Durante, Dirk Hovy, Debora Nozza |
| 2025 | AAAI | SafetyPrompts: A Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety. | Paul Rttger, Fabio Pernisi, Bertie Vidgen, Dirk Hovy |
| 2025 | ACL | The AI Gap: How Socioeconomic Status Affects Language Technology Interactions. | Elisa Bassignana, Amanda Cercas Curry, Dirk Hovy |
| 2025 | ACL | Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions. | Matthias Orlikowski, Jiaxin Pei, Paul Rttger, Philipp Cimiano, David Jurgens, Dirk Hovy |
| 2025 | EMNLP | Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance. | Pedro Henrique Luz de Araujo, Paul Rttger, Dirk Hovy, Benjamin Roth |
| 2025 | EMNLP | Biased Tales: Cultural and Topic Bias in Generating Children's Stories. | Donya Rooein, Vilm Zouhar, Debora Nozza, Dirk Hovy |
| 2025 | EMNLP | Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification. | Chenfei Xiong, Jingwei Ni, Yu Fan, Vilm Zouhar, Donya Rooein, Lorena Calvo-Bartolom, Alexander Miserlis Hoyle, Zhijing Jin, Mrinmaya Sachan, Markus Leippold, Dirk Hovy, Mennatallah El-Assady, Elliott Ash |
| 2024 | ACL | "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models. | Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Rttger, Frauke Kreuter, Dirk Hovy, Barbara Plank |
| 2024 | ACL | Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution. | Flor Miriam Plaza del Arco, Amanda Cercas Curry, Alba Cercas Curry, Gavin Abercrombie, Dirk Hovy |
| 2024 | ACL | Classist Tools: Social Class Correlates with Performance in NLP. | Amanda Cercas Curry, Giuseppe Attanasio, Zeerak Talat, Dirk Hovy |
| 2024 | ACL | Compromesso! Italian Many-Shot Jailbreaks undermine the safety of Large Language Models. | Fabio Pernisi, Dirk Hovy, Paul Rttger |
| 2024 | ACL | Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models. | Paul Rttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schtze, Dirk Hovy |
| 2024 | ACL | Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns. | Antonina Sinelnik, Dirk Hovy |
| 2024 | COLING | Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions. | Flor Miriam Plaza del Arco, Alba Curry, Amanda Cercas Curry, Dirk Hovy |
| 2024 | COLING | Impoverished Language Technology: The Lack of (Social) Class in NLP. | Amanda Cercas Curry, Zeerak Talat, Dirk Hovy |
| 2024 | COLING | DADIT: A Dataset for Demographic Classification of Italian Twitter Users and a Comparison of Prediction Methods. | Lorenzo Lupo, Paul Bose, Mahyar Habibi, Dirk Hovy, Carlo Schwarz |
| 2024 | EACL | Explaining Speech Classification Models via Word-Level Audio Segments and Paralinguistic Features. | Eliana Pastor, Alkis Koudounas, Giuseppe Attanasio, Dirk Hovy, Elena Baralis |
| 2024 | EMNLP | Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models. | Flor Miriam Plaza del Arco, Amanda Cercas Curry, Susanna Paoli, Alba Cercas Curry, Dirk Hovy |
| 2024 | EMNLP | Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps. | Giuseppe Attanasio, Beatrice Savoldi, Dennis Fucci, Dirk Hovy |
| 2024 | NAACL | XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. | Paul Rttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, Dirk Hovy |
| 2023 | ACL | What about "em"? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns. | Anne Lauscher, Debora Nozza, Ehm Miltersen, Archie Crowley, Dirk Hovy |
| 2023 | ACL | The State of Profanity Obfuscation in Natural Language Processing Scientific Publications. | Debora Nozza, Dirk Hovy |
| 2023 | ACL | The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics. | Matthias Orlikowski, Paul Rttger, Philipp Cimiano, Dirk Hovy |
| 2023 | EACL | Can Demographic Factors Improve Text Classification? Revisiting Demographic Adaptation in the Age of Transformers. | Chia-Chien Hung, Anne Lauscher, Dirk Hovy, Simone Paolo Ponzetto, Goran Glavas |
| 2023 | ICWSM | Top-Down Influence? Predicting CEO Personality and Risk Impact from Speech Transcripts. | Kilian Theil, Dirk Hovy, Heiner Stuckenschmidt |
| 2023 | WSDM | Beyond Digital "Echo Chambers": The Role of Viewpoint Diversity in Political Discussion. | Rishav Hada, Amir Ebrahimi Fard, Sarah Shugars, Federico Bianchi, Patrcia G. C. Rossini, Dirk Hovy, Rebekah Tromble, Nava Tintarev |
| 2022 | ACL | Entropy-based Attention Regularization Frees Unintended Bias Mitigation from Lists. | Giuseppe Attanasio, Debora Nozza, Dirk Hovy, Elena Baralis |
| 2022 | ACL | SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems. | Emily Dinan, Gavin Abercrombie, A. Stevie Bergman, Shannon L. Spruit, Dirk Hovy, Y-Lan Boureau, Verena Rieser |
| 2022 | ACL | Hard and Soft Evaluation of NLP models with BOOtSTrap SAmpling - BooStSa. | Tommaso Fornaciari, Alexandra Uma, Massimo Poesio, Dirk Hovy |
| 2022 | COLING | Welcome to the Modern World of Pronouns: Identity-Inclusive Natural Language Processing beyond Gender. | Anne Lauscher, Archie Crowley, Dirk Hovy |
| 2022 | EMNLP | Twitter-Demographer: A Flow-based Tool to Enrich Twitter Data. | Federico Bianchi, Vincenzo Cutrona, Dirk Hovy |
| 2022 | EMNLP | "It's Not Just Hate": A Multi-Dimensional Perspective on Detecting Harmful Speech Online. | Federico Bianchi, Stefanie Anja Hills, Patrcia G. C. Rossini, Dirk Hovy, Rebekah Tromble, Nava Tintarev |
| 2022 | EMNLP | Bridging Fairness and Environmental Sustainability in Natural Language Processing. | Marius Hessenthaler, Emma Strubell, Dirk Hovy, Anne Lauscher |
| 2022 | EMNLP | SocioProbe: What, When, and Where Language Models Learn about Sociodemographics. | Anne Lauscher, Federico Bianchi, Samuel R. Bowman, Dirk Hovy |
| 2022 | EMNLP | Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages. | Paul Rttger, Debora Nozza, Federico Bianchi, Dirk Hovy |
| 2022 | NAACL | Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks. | Paul Rttger, Bertie Vidgen, Dirk Hovy, Janet B. Pierrehumbert |
| 2022 | SIGdial | Guiding the Release of Safer E2E Conversational AI through Value Sensitive Design. | A. Stevie Bergman, Gavin Abercrombie, Shannon L. Spruit, Dirk Hovy, Emily Dinan, Y-Lan Boureau, Verena Rieser |
| 2021 | ACL | On the Gap between Adoption and Understanding in NLP. | Federico Bianchi, Dirk Hovy |
| 2021 | ACL | Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. | Federico Bianchi, Silvia Terragni, Dirk Hovy |
| 2021 | ACL | "We will Reduce Taxes" - Identifying Election Pledges with Language Models. | Tommaso Fornaciari, Dirk Hovy, Elin Naurin, Julia Runeson, Robert Thomson, Pankaj Adhikari |
| 2021 | EACL | Cross-lingual Contextualized Topic Models with Zero-shot Learning. | Federico Bianchi, Silvia Terragni, Dirk Hovy, Debora Nozza, Elisabetta Fersini |
| 2021 | EACL | BERTective: Language Models and Contextual Information for Deception Detection. | Tommaso Fornaciari, Federico Bianchi, Massimo Poesio, Dirk Hovy |
| 2021 | NAACL | Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task Learning. | Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, Massimo Poesio |
| 2021 | NAACL | The Importance of Modeling Social Factors of Language: Theory and Practice. | Dirk Hovy, Diyi Yang |
| 2021 | NAACL | HONEST: Measuring Hurtful Sentence Completion in Language Models. | Debora Nozza, Federico Bianchi, Dirk Hovy |
| 2020 | ACL | Integrating Ethics into the NLP Curriculum. | Emily M. Bender, Dirk Hovy, Alexandra Schofield |
| 2020 | ACL | "You Sound Just Like Your Father" Commercial Machine Translation Systems Include Stylistic Biases. | Dirk Hovy, Federico Bianchi, Tommaso Fornaciari |
| 2020 | ACL | Predictive Biases in Natural Language Processing Models: A Conceptual Framework and Overview. | Deven Shah, H. Andrew Schwartz, Dirk Hovy |
| 2020 | EMNLP | Helpful or Hierarchical? Predicting the Communicative Strategies of Chat Participants, and their Impact on Success. | Farzana Rashid, Tommaso Fornaciari, Dirk Hovy, Eduardo Blanco, Fernando Vega-Redondo |
| 2020 | HCOMP | A Case for Soft Loss Functions. | Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio |
| 2018 | EMNLP | Improving Author Attribute Prediction by Retrofitting Linguistic Representations with Homophily. | Dirk Hovy, Tommaso Fornaciari |
| 2018 | EMNLP | Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting. | Dirk Hovy, Christoph Purschke |
| 2018 | EMNLP | Predicting News Headline Popularity with Syntactic and Semantic Knowledge Using Multi-Task Learning. | Sotiris Lamprinidis, Daniel Hardt, Dirk Hovy |
| 2017 | EACL | Multitask Learning for Mental Health Conditions with Limited Social Media Data. | Adrian Benton, Margaret Mitchell, Dirk Hovy |
| 2017 | EMNLP | End-to-End Information Extraction without Token-Level Supervision. | Rasmus Berg Palm, Dirk Hovy, Florian Laws, Ole Winther |
| 2016 | ACL | Putting Sarcasm Detection into Context: The Effects of Class Imbalance and Manual Labelling on Supervised Machine Classification of Twitter Conversations. | Gavin Abercrombie, Dirk Hovy |
| 2016 | ACL | The Enemy in Your Own Camp: How Well Can We Detect Statistically-Generated Fake Reviews - An Adversarial Study. | Dirk Hovy |
| 2016 | ACL | The Social Impact of Natural Language Processing. | Dirk Hovy, Shannon L. Spruit |
| 2016 | LREC | Exploring Language Variation Across Europe - A Web-based Tool for Computational Sociolinguistics. | Dirk Hovy, Anders Johannsen |
| 2016 | NAACL | Learning a POS tagger for AAVE-like language. | Anna Jrgensen, Dirk Hovy, Anders Sgaard |
| 2016 | NAACL | Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter. | Zeerak Waseem, Dirk Hovy |
| 2015 | ACL | If all you have is a bit of the Bible: Learning POS taggers for truly low-resource languages. | Zeljko Agic, Dirk Hovy, Anders Sgaard |
| 2015 | ACL | Demographic Factors Improve Classification Performance. | Dirk Hovy |
| 2015 | ACL | Tagging Performance Correlates with Author Age. | Dirk Hovy, Anders Sgaard |
| 2015 | CoNLL | Cross-lingual syntactic variation over age and gender. | Anders Johannsen, Dirk Hovy, Anders Sgaard |
| 2015 | EMNLP | The Rating Game: Sentiment Rating Reproducibility from Text. | Lasse Borgholt, Peter Simonsen, Dirk Hovy |
| 2015 | NAACL | Mining for unambiguous instances to adapt part-of-speech taggers to new domains. | Dirk Hovy, Barbara Plank, Hctor Martnez Alonso, Anders Sgaard |
| 2015 | WWW | User Review Sites as a Resource for Large-Scale Sociolinguistic Studies. | Dirk Hovy, Anders Johannsen, Anders Sgaard |
| 2014 | ACL | How Well can We Learn Interpretable Entity Types from Text? | Dirk Hovy |
| 2014 | ACL | Experiments with crowdsourced re-annotation of a POS tagging data set. | Dirk Hovy, Barbara Plank, Anders Sgaard |
| 2014 | ACL | Linguistically debatable or just plain wrong? | Barbara Plank, Dirk Hovy, Anders Sgaard |
| 2014 | COLING | Adapting taggers to Twitter with not-so-distant supervision. | Barbara Plank, Dirk Hovy, Ryan T. McDonald, Anders Sgaard |
| 2014 | COLING | Selection Bias, Label Bias, and Bias in Ground Truth. | Anders Sgaard, Barbara Plank, Dirk Hovy |
| 2014 | CoNLL | What's in a p-value in NLP? | Anders Sgaard, Anders Johannsen, Barbara Plank, Dirk Hovy, Hctor Martnez Alonso |
| 2014 | EACL | Learning part-of-speech taggers with inter-annotator agreement loss. | Barbara Plank, Dirk Hovy, Anders Sgaard |
| 2014 | LREC | Crowdsourcing and annotating NER for Twitter #drift. | Hege Fromreide, Dirk Hovy, Anders Sgaard |
| 2014 | LREC | When POS data sets don't add up: Combatting sample bias. | Dirk Hovy, Barbara Plank, Anders Sgaard |
| 2014 | LREC | Augmenting English Adjective Senses with Supersenses. | Yulia Tsvetkov, Nathan Schneider, Dirk Hovy, Archna Bhatia, Manaal Faruqui, Chris Dyer |
| 2013 | EMNLP | A Walk-Based Semantically Enriched Tree Kernel Over Distributed Word Representations. | Shashank Srivastava, Dirk Hovy, Eduard H. Hovy |
| 2013 | Interspeech | Analysis and modeling of "focus" in context. | Dirk Hovy, Gopala Krishna Anumanchipalli, Alok Parlikar, Caroline Vaughn, Adam C. Lammert, Eduard H. Hovy, Alan W. Black |
| 2013 | NAACL | Learning Whom to Trust with MACE. | Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, Eduard H. Hovy |
| 2013 | WWW | Solving electrical networks to incorporate supervision in random walks. | Mrinmaya Sachan, Dirk Hovy, Eduard H. Hovy |
| 2012 | EACL | When Did that Happen? - Linking Events and Relations to Timestamps. | Dirk Hovy, James Fan, Alfio Massimiliano Gliozzo, Siddharth Patwardhan, Christopher A. Welty |
| 2011 | ACL | Models and Training for Unsupervised Preposition Sense Disambiguation. | Dirk Hovy, Ashish Vaswani, Stephen Tratz, David Chiang, Eduard H. Hovy |
| 2011 | ACL | Unsupervised Discovery of Domain-Specific Knowledge from Text. | Dirk Hovy, Chunliang Zhang, Eduard H. Hovy, Anselmo Peas |
| 2010 | COLING | What's in a Preposition? Dimensions of Sense Disambiguation for an Interesting Word Class. | Dirk Hovy, Stephen Tratz, Eduard H. Hovy |
| 2009 | NAACL | Disambiguation of Preposition Sense Using Linguistically Motivated Features. | Stephen Tratz, Dirk Hovy |