| 2024 | A Hypothesis-Driven Framework for the Analysis of Self-Rationalising Models. | Marc Braun, Jenny Kunz |
| 2024 | A Thesis Proposal ClaimInspector Framework: A Hybrid Approach to Data Annotation using Fact-Checked Claims and LLMs. | Basak Bozkurt |
| 2024 | SENSE-LM : A Synergy between a Language Model and Sensorimotor Representations for Auditory and Olfactory Information Extraction. | Cdric Boscher, Christine Largeron, Vronique Eglin, Eld Egyed-Zsigmond |
| 2024 | Centering the Speech Community. | Steven Bird, Dean Yibarbuk |
| 2024 | AnaDE1.0: A Novel Data Set for Benchmarking Analogy Detection and Extraction. | Bhavya, Shradha Sehgal, Jinjun Xiong, ChengXiang Zhai |
| 2024 | Tsetlin Machine Embedding: Representing Words Using Logical Expressions. | Bimal Bhattarai, Ole-Christoffer Granmo, Lei Jiao, Rohan Kumar Yadav, Jivitesh Sharma |
| 2024 | Rainbow - A Benchmark for Systematic Testing of How Sensitive Visio-Linguistic Models are to Color Naming. | Marie Bexte, Andrea Horbach, Torsten Zesch |
| 2024 | LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text. | Dor Bernsohn, Gil Semo, Yaron Vazana, Gila Hayat, Ben Hagag, Joel Niklaus, Rohit Saha, Kyryl Truskovskyi |
| 2024 | Cross-lingual Editing in Multilingual Language Models. | Himanshu Beniwal, Kowsik Nandagopan D, Mayank Singh |
| 2024 | PRILoRA: Pruned and Rank-Increasing Low-Rank Adaptation. | Nadav Benedek, Lior Wolf |
| 2024 | Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting. | Tilman Beck, Hendrik Schuff, Anne Lauscher, Iryna Gurevych |
| 2024 | Multilingual Gradient Word-Order Typology from Universal Dependencies. | Emi Baylor, Esther Ploeger, Johannes Bjerva |
| 2024 | Testing the Depth of ChatGPT's Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5's Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking. | David Bayani |
| 2024 | Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification. | Luke Bates, Iryna Gurevych |
| 2024 | EXPLORER: Exploration-guided Reasoning for Textual Reinforcement Learning. | Kinjal Basu, Keerthiram Murugesan, Subhajit Chaudhury, Murray Campbell, Kartik Talamadupula, Tim Klinger |
| 2024 | Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs. | Simone Balloccu, Patrcia Schmidtov, Mateusz Lango, Ondrej Dusek |
| 2024 | FAIR: Filtering of Automatically Induced Rules. | Divya Jyoti Bajpai, Ayush Maheshwari, Manjesh Kumar Hanawal, Ganesh Ramakrishnan |
| 2024 | Threat Behavior Textual Search by Attention Graph Isomorphism. | Chanwoo Bae, Guanhong Tao, Zhuo Zhang, Xiangyu Zhang |
| 2024 | Interpreting Predictive Probabilities: Model Confidence or Human Label Variation? | Joris Baan, Raquel Fernndez, Barbara Plank, Wilker Aziz |
| 2024 | Topic-guided Example Selection for Domain Adaptation in LLM-based Machine Translation. | Seth Aycock, Rachel Bawden |
| 2024 | Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties. | Ekaterina Artemova, Verena Blaschke, Barbara Plank |
| 2024 | Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models. | Erik Arakelyan, Zhaoqi Liu, Isabelle Augenstein |
| 2024 | Frchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels. | Negar Arabzadeh, Charles L. A. Clarke |
| 2024 | Capturing the Relationship Between Sentence Triplets for LLM and Human-Generated Texts to Enhance Sentence Embeddings. | Na Min An, Sania Waheed, James Thorne |
| 2024 | Dynamic Masking Rate Schedules for MLM Pretraining. | Zachary Ankner, Naomi Saphra, Davis W. Blalock, Jonathan Frankle, Matthew L. Leavitt |