| 2025 | Transformative or Conservative? Conservation laws for ResNets and Transformers. | Sibylle Marcotte, Rmi Gribonval, Gabriel Peyr |
| 2025 | Position: Algebra Unveils Deep Learning - An Invitation to Neuroalgebraic Geometry. | Giovanni Luca Marchetti, Vahid Shahverdi, Stefano Mereta, Matthew Trager, Kathln Kohn |
| 2025 | Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity. | Arto Maranjyan, Alexander Tyurin, Peter Richtrik |
| 2025 | ATA: Adaptive Task Allocation for Efficient Resource Management in Distributed Machine Learning. | Arto Maranjyan, El Mehdi Saad, Peter Richtrik, Francesco Orabona |
| 2025 | COExpander: Adaptive Solution Expansion for Combinatorial Optimization. | Jiale Ma, Wenzheng Pan, Yang Li, Junchi Yan |
| 2025 | PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling. | Avery Ma, Yangchen Pan, Amir-massoud Farahmand |
| 2025 | Principled Algorithms for Optimizing Generalized Metrics in Binary Classification. | Anqi Mao, Mehryar Mohri, Yutao Zhong |
| 2025 | Mastering Multiple-Expert Routing: Realizable H-Consistency and Strong Guarantees for Learning to Defer. | Anqi Mao, Mehryar Mohri, Yutao Zhong |
| 2025 | Hypo3D: Exploring Hypothetical Reasoning in 3D. | Ye Mao, Weixun Luo, Junpeng Jing, Anlan Qiu, Krystian Mikolajczyk |
| 2025 | CTBench: A Library and Benchmark for Certified Training. | Yuhao Mao, Stefan Balauca, Martin T. Vechev |
| 2025 | TimePro: Efficient Multivariate Long-term Time Series Forecasting with Variable- and Time-Aware Hyper-state. | Xiaowen Ma, Zhen-Liang Ni, Shuai Xiao, Xinghao Chen |
| 2025 | Efficient Diffusion Models for Symmetric Manifolds. | Oren Mangoubi, Neil He, Nisheeth K. Vishnoi |
| 2025 | Scaffold with Stochastic Gradients: New Analysis with Linear Speed-Up. | Paul Mangold, Alain Oliviero Durmus, Aymeric Dieuleveut, Eric Moulines |
| 2025 | Learning Latent Graph Structures and their Uncertainty. | Alessandro Manenti, Daniele Zambon, Cesare Alippi |
| 2025 | Potemkin Understanding in Large Language Models. | Marina Mancoridis, Bec Weeks, Keyon Vafa, Sendhil Mullainathan |
| 2025 | Aggregation of Dependent Expert Distributions in Multimodal Variational Autoencoders. | Rogelio Andrade Mancisidor, Robert Jenssen, Shujian Yu, Michael Kampffmeyer |
| 2025 | Learning from Loss Landscape: Generalizable Mixed-Precision Quantization via Adaptive Sharpness-Aware Gradient Aligning. | Lianbo Ma, Jianlun Ma, Yuee Zhou, Guoyang Xie, Qiang He, Zhichao Lu |
| 2025 | Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning. | Guozheng Ma, Lu Li, Zilin Wang, Li Shen, Pierre-Luc Bacon, Dacheng Tao |
| 2025 | Catching Two Birds with One Stone: Reward Shaping with Dual Random Networks for Balancing Exploration and Exploitation. | Haozhe Ma, Fangling Li, Jing Yu Lim, Zhengding Luo, Thanh Vinh Vo, Tze-Yun Leong |
| 2025 | Emergence in non-neural models: grokking modular arithmetic via average gradient outer product. | Neil Mallinar, Daniel Beaglehole, Libin Zhu, Adityanarayanan Radhakrishnan, Parthe Pandit, Mikhail Belkin |
| 2025 | Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard Problems. | Mikolaj Malkinski, Szymon Pawlonka, Jacek Mandziuk |
| 2025 | Craftium: Bridging Flexibility and Efficiency for Rich 3D Single- and Multi-Agent Environments. | Mikel Malagn, Josu Ceberio, Jos Antonio Lozano |
| 2025 | PTTA: Purifying Malicious Samples for Test-Time Model Adaptation. | Jing Ma, Hanlin Li, Xiang Xiang |
| 2025 | TANGO: Clustering with Typicality-Aware Nonlocal Mode-Seeking and Graph-Cut Optimization. | Haowen Ma, Zhiguo Long, Hua Meng |
| 2025 | Hypothesis Testing for Generalized Thurstone Models. | Anuran Makur, Japneet Singh |