| 2026 | Comparing Human and Large Language Model Interpretation of Implicit Information. | Antonio De Santis, Tommaso Bonetti, Andrea Tocchetti, Marco Brambilla |
| 2026 | Generalization or Memorization? Multi-Agent vs. Baseline LLMs and AutoML Models for Tabular Classification. | Aida Sanatizadeh, Sorouralsadat Fatemi, Reza Mousavi, Ahmed Abbasi |
| 2026 | Representing Lean Proofs as Trajectories in Latent Space. | Elisaveta Samoylov, Soroush Vosoughi |
| 2026 | MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication. | Sraavya Sambara, Yuan Pu, Ayman Ali, Vishala Mishra, Lionel Wong, Monica Agrawal |
| 2026 | Simple Agents, Biased Judges: Efficient Multi-Party Dialogue Generation & The Evaluation Gap. | Kunal Samanta, Faisal Tareque Shohan, Amine Trabelsi, Richard Khoury |
| 2026 | Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models. | Mehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain, Wasi Uddin Ahmad, Somshubra Majumdar, Vahid Noroozi, Boris Ginsburg |
| 2026 | MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following. | Mohammad Mahdi Salmani-Zarchi, Zahra Rahimi, Heshaam Faili, Mohammad Javad Dousti |
| 2026 | Expert Calibration Lens for Pruning Mixture of Experts. | Luis Frentzen Salim, Chia-Chun Wu, Tran Van Nhiem, Lun-Wei Ku, Yung-Hui Li |
| 2026 | Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs. | Israfel Salazar, Desmond Elliott, Yova Kementchedjhieva |
| 2026 | Through the Looking Glass of Multilingual AI: Contrasting Language- and Name Script-Dependent Ethnic Hierarchies in GPT and DeepSeek. | Annabella Sakunkoo, Jonathan Sakunkoo |
| 2026 | Can Large Language Models Infer Causal Relationships from Real-World Text? | Ryan Saklad, Aman Chadha, Oleg V. Pavlov, Raha Moraffah |
| 2026 | Thinking Like a Botanist: Challenging Multimodal Language Models with Intent Driven Chain-of-Inquiry. | Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin, Hasan Muhammad Abdullah, Md. Mehedi Hasan, Mohammad Zabed Hossain, Shifat E. Arman |
| 2026 | HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences. | Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe |
| 2026 | Towards Understanding the Robustness of Sparse Autoencoders. | Ahson Saiyed, Sabrina Sadiekh, Chirag Agarwal |
| 2026 | Convergent Demographic Utility Hierarchies: Geometry of Intersectional Values in LLMs. | Pravish Sainath |
| 2026 | Do Emotions Influence Moral Judgment in Large Language Models? | Mohammad Saim, Tianyu Jiang |
| 2026 | Reward Modeling for Scientific Writing Evaluation. | Furkan Sahinu, Subhabrata Dutta, Iryna Gurevych |
| 2026 | Reference-Free Schema Generation for Literature Review Tables via Multi-Faceted Rewards. | Sinjoy Saha, Suman Saha, Mahfuza Farooque, Wenpeng Yin |
| 2026 | Zero-Shot Multimodal Retrieval with Multi-Scale Contextual Representations. | Sourajit Saha, Tejas Gokhale |
| 2026 | SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks. | Mohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Oroojlooy, Graham Horwood, Dan Roth |
| 2026 | Evaluating Perspectival Biases in Cross-Modal Retrieval. | Teerapol Saengsukhiran, Peerawat Chomphooyod, Narabodee Rodjananant, Chompakorn Chaksangchaichot, Patawee Prakrankamanant, Witthawin Sripheanpol, Pak Lovichit, Sarana Nutanong, Ekapol Chuangsuwanich |
| 2026 | FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments. | Amir Saeidi, Venkatesh Mishra, Souradeep Mukhopadhyay, Gaowen Liu, Ali Payani, Jayanth Srinivasa, Chitta Baral |
| 2026 | Instruction-Guided Poetry Generation in Arabic and Its Dialects. | Abdelrahman Boda Sadallah, Kareem Ashraf Elozeiri, Mervat Abassy, Rania Elbadry, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Fajri Koto |
| 2026 | What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment? | Koki Ryu, Hitomi Yanaka |
| 2026 | Feeling Right vs. Being Right: How AI Sycophancy Affects Value-Laden Deliberation. | Jeongwoo Ryu, Soomin Kim, Jinsu Eun, Kyusik Kim, Changhoon Oh, Bongwon Suh |