| 2026 | Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation. | Srujan P. Mule, Aniketh Garikaparthi, Manasi Patwardhan |
| 2026 | Credal Concept Bottleneck Models for Epistemic-Aleatoric Uncertainty Decomposition. | Tanmoy Mukherjee, Thomas Bailleux, Pierre Marquis, Zied Bouraoui |
| 2026 | Stress Testing Factual Consistency Metrics for Long-Document Summarization. | Zain Muhammad Mujahid, Dustin Wright, Isabelle Augenstein |
| 2026 | Counterspeech Generation using Small Language Models. | Abubakar Sadiq Muhammad, Simona Frenda, Gavin Abercrombie |
| 2026 | From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts? | Aaron Mueller, Andrew Lee, Shruti Joshi, Ekdeep Singh Lubana, Dhanya Sridhar, Patrik Reizinger |
| 2026 | DIXITWORLD: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay. | Yunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu, Baixuan Xu, Yauwai Yim, Chunkit Chan, Jiaxin Bai, Yangqiu Song |
| 2026 | Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring. | Jamshid Mozafari, Bhawna Piryani, Adam Jatowt |
| 2026 | Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence. | Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib, Ramez Kouzy, Wei Xu, Byron C. Wallace, Junyi Jessy Li |
| 2026 | J-Shuwa: A Large-Scale Web-Collected Japanese Sign Language-Japanese Parallel Corpus. | Junwen Mo, MinhDuc Vo, Noriki Nishida, Shin'ichi Satoh, Hideki Nakayama |
| 2026 | ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback. | Yutao Mou, Zhangchi Xue, Lijun Li, Peiyang Liu, Shikun Zhang, Wei Ye, Jing Shao |
| 2026 | Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models. | Basel Mousi, Fahim Dalvi, Shammur Absar Chowdhury, Firoj Alam, Nadir Durrani |
| 2026 | AdabNER: Arabic Digital Archive Books with Nested Entity Recognition. | Aya Mourad, Mustafa Jarrar |
| 2026 | Thinking Twice Makes Large Language Models Safer and More Helpful. | Yutao Mou, Yuxiao Luo, Shikun Zhang, Wei Ye |
| 2026 | Trajectory Signatures of Deception in Large Language Models. | Viraaji Mothukuri, Reza M. Parizi |
| 2026 | Task Assignment meets Annotator Modeling: Human-LLM Collaborative Annotation with Constraints. | Kei Moriyama, Kouta Nakayama, Yukino Baba |
| 2026 | Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG). | Junki Mori, Kazuya Kakizaki, Taiki Miyagawa, Jun Sakuma |
| 2026 | Benchmarking Deflection and Hallucination in Large Vision-Language Models. | Nicholas Moratelli, Christopher Davis, Leonardo F. R. Ribeiro, Bill Byrne, Gonzalo Iglesias |
| 2026 | ToMMeR - Efficient Entity Mention Detection from Large Language Models. | Victor Morand, Nadi Tomeh, Josiane Mothe, Benjamin Piwowarski |
| 2026 | FlowHN: Adaptive Token Routing for Efficient Parallel Hybrid Networks. | Mohammad Mahdi Moradi, Walid Ahmed, Shuangyue Wen, Sudhir Mudur, Weiwei Zhang, Yang Liu |
| 2026 | Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs. | Guy Mor-Lan, Omer Goldman, Matan Eyal, Adi Mayrav Gilady, Sivan Eiger, Idan Szpektor, Avinatan Hassidim, Yossi Matias, Reut Tsarfaty |
| 2026 | NeedleChain: Measuring Intact Context Comprehension Capability of Large Language Models. | Hyeonseok Moon, Heuiseok Lim |
| 2026 | Towards A Scanpath-Conditioned Surprisal Theory: Modeling Reader Information States. | Michael Mooney, Edmond S. L. Ho |
| 2026 | Lightweight and Faithful Visual Condition Checking in Behavior Trees via Expert-Regularized Reinforcement Learning. | Hyosik Moon, Eldan Cohen |
| 2026 | From Fluent to Useful: Generative AI That Models Purpose, Audience, and Presenter for Scientific Communication. | Ishani Mondal |
| 2026 | RedCoder: Automated Multi-Turn Red Teaming for Code LLMs. | Wenjie Jacky Mo, Qin Liu, Xiaofei Wen, Dongwon Jung, Hadi Askari, Wenxuan Zhou, Zhe Zhao, Muhao Chen |