| 2026 | Evaluation Pitfalls and Challenges in Multimedia Event Extraction. | Philipp Seeberger, Steffen Freisinger, Tobias Bocklet, Korbinian Riedhammer |
| 2026 | Whose Facts Win? LLM Source Preferences under Knowledge Conflicts. | Jakob Schuster, Vagrant Gautam, Katja Markert |
| 2026 | Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias. | Elias Schuhmacher, Andrianos Michail, Juri Opitz, Rico Sennrich, Simon Clematide |
| 2026 | Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models. | Tobias Schreieder, Tim Schopf, Michael Frber |
| 2026 | Compact Example-Based Explanations for Language Models. | Loris Schoenegger, Benjamin Roth |
| 2026 | Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains. | Finn Schmidt, Jan Philip Wahle, Terry Ruas, Bela Gipp |
| 2026 | Probabilistic Depression Detection from Textual Time Series. | Fabian Schmidt, Seyedehmoniba Ravan, Vladimir Vlassov |
| 2026 | Similarity-Distance-Magnitude Activations. | Allen Schmaltz |
| 2026 | CLARO: Controlled Attribute-Driven Reasoning Optimization for Efficient Chain-of-Thought. | Oded Schlesinger, Young Kyung Kim, J. Matas Di Martino, Guillermo Sapiro |
| 2026 | Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning. | Max Schaffelder, Albert Gatt |
| 2026 | Simulated Students in Tutoring Dialogues: Substance or Illusion? | Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew Lan |
| 2026 | Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains. | Md. Faiyaz Abdullah Sayeedi, Subhey Sadi Rahman, Md. Mahbub Alam, Md. Adnanul Islam, Jannatul Ferdous Deepti, Tasnim Mohiuddin, Md Mofijul Islam, Swakkhar Shatabda |
| 2026 | Actionable Interpretability for Churn Classification: A Text Bottleneck Model Case Study at a Major Telecom Provider. | Adrian Sauter, Vera Neplenbroek, Georgios Vlassopoulos, Gianluigi Bardelloni |
| 2026 | Evaluating Reasoning Models for Queries with Presuppositions. | Rose Sathyanathan, Kinshuk Vasisht, Danish Pruthi |
| 2026 | Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions. | Ignacio Sastre, Aiala Ros |
| 2026 | CLARITY: A Framework and Benchmark for Conversational Language Ambiguity and Unanswerability in Interactive NL2SQL Systems. | Tabinda Sarwar, Farhad Moghimifar, Cong Duy Vu Hoang, Xiaoxiao Ma, Shawn Chang Xu, Fahimeh Saleh, Poorya Zaremoodi, Avirup Sil, Katrin Kirchhoff |
| 2026 | Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks. | Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu, Vaidehi Patil |
| 2026 | Multi-Constraint State Tracking with Negation: A Diagnostic Benchmark for LLM World Modeling. | Ayan Sar, Pranav Singh Puri, Sumit Aich, Anurag Kaushish, Tanupriya Choudhury, Ajith Abraham |
| 2026 | Digitizing Nepal's Written Heritage: A Comprehensive HTR Pipeline for Old Nepali Manuscripts. | Anjali Sarawgi, Esteban Garces Arias, Christof Zotter |
| 2026 | AURORA: Neuro-Symbolic Continual Indexing for Evolving RAG Systems. | Manoj Saravanan, Rohit Kumar Salla, Ramya Manasa Amancherla |
| 2026 | Can Small LLMs Learn a Robust Theory of Mind via RLVR? Investigating Generalization through the False-Belief Task. | Sneheel Sarangi, Hanan Salam |
| 2026 | Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning. | Raman Saparkhan, Majd Hawasly, Md. Rizwan Parvez, Mohammad Raza |
| 2026 | Large Language Models Are Overconfident in Their Own Responses. | Mario Sanz-Guerrero, Manuel Mager, Katharina von der Wense |
| 2026 | Mask-to-Correct⁺: Leveraging Retriever Diversity for Masking-guided Faithful Fact Correction. | Payel Santra, Lavisha Sharma, Madhusudan Ghosh, Partha Basuchowdhuri |
| 2026 | From Naturalness to Norms: Interactional Cultural Competence for SpeechLMs. | T. Y. S. S. Santosh |