| 2025 | ProofTeller: Exposing recency bias in LLM reasoning and its side effects on communication. | Mayank Jobanputra, Alisa Kovtunova, Brisca Balthes, Fedor Grigoryevich Pogulskiy, Yifan Wang, Stefan Borgwardt, Vera Demberg |
| 2025 | Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing. | Jiabao Ji, Bairu Hou, Alexander Robey, George J. Pappas, Hamed Hassani, Yang Zhang, Eric Wong, Shiyu Chang |
| 2025 | Topology-Aware Gated Graph Neural Network for Social Bot Detection. | Pi Jiebin, Yantuan Xian, Yuxin Huang, Yan Xiang, Ran Song, Zhengtao Yu |
| 2025 | HARBOR: Exploring Persona Dynamics in Multi-Agent Competition. | Kenan Jiang, Li Xiong, Fei Liu |
| 2025 | What Would You Ask When You First Saw a²+b²=c²? Evaluating LLM on Curiosity-Driven Question Generation. | Shashidhar Reddy Javaji, Zining Zhu |
| 2025 | Can AI Validate Science? Benchmarking LLMs on Claim →Evidence Reasoning in AI Papers. | Shashidhar Reddy Javaji, Yupeng Cao, Haohang Li, Yangyang Yu, Nikhil Muralidhar, Zining Zhu |
| 2025 | Teaching Sarcasm: Few-Shot Multimodal Sarcasm Detection via Distillation to a Parameter-Efficient Student. | Soumyadeep Jana, Sanasam Ranbir Singh |
| 2025 | Modeling Contextual Passage Utility for Multihop Question Answering. | Akriti Jain, Aparna Garimella |
| 2025 | INTERCHART: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information. | Anirudh Iyengar Kaniyar Narayana Iyengar, Srija Mukhopadhyay, Adnan Qidwai, Shubhankar Singh, Dan Roth, Vivek Gupta |
| 2025 | Fine-grained Confidence Estimation for Spurious Correctness Detection in Large Language Models. | Ai Ishii, Naoya Inoue, Hisami Suzuki, Satoshi Sekine |
| 2025 | Multi-Agent Cross-Lingual Veracity Assessment for Explainable Fake News Detection. | Bassamtiano Renaufalgi Irnawan, Yoshimi Suzuki, Noriko Tomuro, Fumiyo Fukumoto |
| 2025 | The Visual Counter Turing Test (VCT²): A Benchmark for Evaluating AI-Generated Image Detection and the Visual AI Index (V_AI). | Nasrin Imanpour, Abhilekh Borah, Shashwat Bajpai, Subhankar Ghosh, Sainath Reddy Sankepally, Hasnat Md Abdullah, Nishoak Kosaraju, Shreyas Dixit, Ashhar Aziz, Shwetangshu Biswas, Vinija Jain, Aman Chadha, Song Wang, Amit P. Sheth, Amitava Das |
| 2025 | Ability Transfer Through Language Mixing. | Petr Hyner, Jan Mrgala, Jan Hula |
| 2025 | Simple Yet Effective: Extracting Private Data Across Clients in Federated Fine-Tuning of Large Language Models. | Yingqi Hu, Zhuo Zhang, Jingyuan Zhang, Jinghua Wang, Qifan Wang, Lizhen Qu, Zenglin Xu |
| 2025 | Characterizing Mamba's Selective Memory using Auto-Encoders. | Tamanna Hossain, Robert L. Logan IV, Ganesh Jagadeesan, Sameer Singh, Joel R. Tetreault, Alejandro Jaimes |
| 2025 | SciHallu: A Multi-Granularity Hallucination Detection Dataset for Scientific Writing. | Adiba Ibnat Hossain, Sagnik Ray Choudhury, Hamed Alhoori |
| 2025 | Intrinsic Linguistic Bias in Formal vs. Informal Bengali Pragmatics with Progressive Context Inflation. | Md. Tanzib Hosain, Md. Kishor Morol |
| 2025 | Exploring Working Memory Capacity in LLMs: From Stressors to Human-Inspired Strategies. | Eunjin Hong, Sumin Cho, Juae Kim |
| 2025 | Large Temporal Models: Unlocking Temporal Understanding in LLMs for Temporal Relation Classification. | Omri Homburger, Kfir Bar |
| 2025 | Investigating Feasibility of Large Language Model Agent Collaboration in Minecraft and Comparison with Human-Human Collaboration. | Yuki Hirota, Ryuichiro Higashinaka |
| 2025 | VAGUE-Gate: Plug-and-Play Local-Privacy Shield for Retrieval-Augmented Generation. | Arshia Hemmat, Matin Moqadas, Ali Mamanpoosh, Amirmasoud Rismanchian, Afsaneh Fatemi |
| 2025 | Not Just a Piece of Cake: Cross-Lingual Fine-Tuning for Idiom Identification. | Ofri Hefetz, Kai Golan Hashiloni, Alon Mannor, Kfir Bar |
| 2025 | DharmaBench: Evaluating Language Models on Buddhist Texts in Sanskrit and Tibetan. | Kai Golan Hashiloni, Shay Cohen, Asaf Shina, Jingyi Yang, Orr Meir Zwebner, Nicola Bajetta, Guy Bilitski, Rebecca Sundn, Guy Maduel, Ryan Conlon, Ari Barzilai, Daniel Mass, Shanshan Jia, Aviv Naaman, Sonam Choden, Sonam Jamtsho, Yadi Qu, Harunaga Isaacson, Dorji Wangchuk, Shai Fine, Orna Almogi, Kfir Bar |
| 2025 | "Whose Side Are You On?" Estimating Ideology of Political and News Content Using Large Language Models and Few-shot Demonstration Selection. | Muhammad Haroon, Magdalena Wojcieszak, Anshuman Chhabra |
| 2025 | An Analysis of the Impact of Problem Paraphrasing on LLM-Based Mathematical Problem Solving. | Yerim Han, Hyein Seo, Hyuk Namgoong, Sangkeun Jung |