| 2025 | MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools. | Nishant Subramani, Jason Eisner, Justin Svegliato, Benjamin Van Durme, Yu Su, Sam Thomson |
| 2025 | HARP: Hesitation-Aware Reframing in Transformer Inference Pass. | Romain Stora, Seung-won Hwang |
| 2025 | Teaching Models to Balance Resisting and Accepting Persuasion. | Elias Stengel-Eskin, Peter Hase, Mohit Bansal |
| 2025 | MonoTODia: Translating Monologue Requests to Task-Oriented Dialogues. | Sebastian Steindl, Ulrich Schfer, Bernd Ludwig |
| 2025 | Dis2Dis: Explaining Ambiguity in Fact-Checking. | Ieva Staliunaite, Andreas Vlachos |
| 2025 | Predicting the Target Word of Game-playing Conversations using a Low-Rank Dialect Adapter for Decoder Models. | Dipankar Srirag, Aditya Joshi, Jacob Eisenstein |
| 2025 | NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals. | Neha Srikanth, Rachel Rudinger |
| 2025 | Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models. | Arvind Krishna Sridhar, Yinyi Guo, Erik Visser |
| 2025 | Adaptive Prompting: Ad-hoc Prompt Composition for Social Bias Detection. | Maximilian Spliethver, Tim Knebler, Fabian Fumagalli, Maximilian Muschalik, Barbara Hammer, Eyke Hllermeier, Henning Wachsmuth |
| 2025 | Chatbot Arena Estimate: towards a generalized performance benchmark for LLM capabilities. | Lucas Spangher, Tianle Li, William F. Arnold, Nick Masiewicki, Xerxes Dotiwalla, Rama Kumar Pasumarthi, Peter Grabowski, Eugene Ie, Daniel Gruhl |
| 2025 | Synthetic Audio Helps for Cognitive State Tasks. | Adil Soubki, John Murzaku, Peter Zeng, Owen Rambow |
| 2025 | Modeling the Differential Prevalence of Online Supportive Interactions in Private Instant Messages of Adolescents. | Ondrej Sotolr, Michal Tkaczyk, Jaromr Plhk, David Smahel |
| 2025 | Hard Emotion Test Evaluation Sets for Language Models. | Tiberiu Sosea, Cornelia Caragea |
| 2025 | CCT-Code: Cross-Consistency Training for Multilingual Clone Detection and Code Search. | Nikita Sorokin, Tikhonov Anton, Dmitry Abulkhanov, Ivan Sedykh, Irina Piontkovskaya, Valentin Malykh |
| 2025 | Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models. | Hyegang Son, Yonglak Son, Changhoon Kim, Young Geun Kim |
| 2025 | KMMLU: Measuring Massive Multitask Language Understanding in Korean. | Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, Stella Biderman |
| 2025 | From Curiosity to Clarity : Exploring the Impact of Consecutive Why-Questions. | Geonyeong Son, Jaeyoung Lee, Misuk Kim |
| 2025 | Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks. | Nikita Soni, Pranav Chitale, Khushboo Singh, Niranjan Balasubramanian, H. Andrew Schwartz |
| 2025 | Learning to Summarize from LLM-generated Feedback. | Hwanjun Song, Taewon Yun, Yuho Lee, Jihwan Oh, Gihun Lee, Jason Cai, Hang Su |
| 2025 | Unified Automated Essay Scoring and Grammatical Error Correction. | Seungwoo Song, Junghun Yuk, ChangSu Choi, Hangyeol Yoo, HyeonSeok Lim, KyungTae Lim, Jungyeul Park |
| 2025 | A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models. | Xiujie Song, Mengyue Wu, Kenny Q. Zhu, Chunhao Zhang, Yanyi Chen |
| 2025 | The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism. | Yifan Song, Guoyin Wang, Sujian Li, Bill Yuchen Lin |
| 2025 | Echoes of Discord: Forecasting Hater Reactions to Counterspeech. | Xiaoying Song, Sharon Lisseth Perez, Xinchen Yu, Eduardo Blanco, Lingzi Hong |
| 2025 | Hazards in Daily Life? Enabling Robots to Proactively Detect and Resolve Anomalies. | Zirui Song, Guangxian Ouyang, Meng Fang, Hongbin Na, Zijing Shi, Zhenhao Chen, Yujie Fu, Zeyu Zhang, Shiyu Jiang, Miao Fang, Ling Chen, Xiuying Chen |
| 2025 | Is a Peeled Apple Still Red? Evaluating LLMs' Ability for Conceptual Combination with Property Type. | Seokwon Song, Taehyun Lee, Jaewoo Ahn, Jae Hyuk Sung, Gunhee Kim |