| 2025 | Towards Adaptable and Intelligible Speech Synthesis in Noisy Environments. | Lubos Marcinek, Jonas Beskow, Joakim Gustafson |
| 2025 | Data-driven approaches to pitch modelling in two Mexican Spanish ethnolects: K-means Clustering & GAMMs. | Gilly Marchini, Jeremy Steffman |
| 2025 | Beat gestures made by human-like avatars affect speech perception. | Matteo Maran, Renske Rtjes, Anna R. E. Schreurs, Hans Rutger Bosker |
| 2025 | Understanding Dementia Speech Alignment with Diffusion-Based Image Generation. | Mansi, Anastasios Lepipas, Dominika C. Woszczyk, Yiying Guan, Soteris Demetriou |
| 2025 | Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models. | Potsawee Manakul, Guangzhi Sun, Warit Sirichotedumrong, Kasima Tharnpipitchai, Kunat Pipatanakul |
| 2025 | STCON NIST SRE24 System: Composite Speaker Recognition Solution for Challenging Scenarios. | Stepan Malykh, Alexander Anikin, Nikita Khmelev, Anastasia Korenevskaya, Anastasia Zorkina, Sergey Novoselov, Vladislav Marchevskiy, Vladimir Volokhov, Andrey Shulipa, Alexander Kozlov, Alexander Melnikov, Vasiliy Galyuk, Timur Pekhovskiy |
| 2025 | SupraDoRAL: Automatic Word Prominence Detection Using Suprasegmental Dependencies of Representations with Acoustic and Linguistic Context. | Jhansi Mallela, Upendra Vishwanath Y. S., Sankara Bharadwaj Rangavajjala, Bhaskar Bhatt, Chiranjeevi Yarra |
| 2025 | Contextual predictability effects on acoustic distinctiveness in read Polish speech. | Zofia Malisz, Jan Foremski, Malgorzata Kul |
| 2025 | DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation. | Prabash Reddy Male, Swayambhu Nath Ray, Harish Arsikere, Akshat Jaiswal, Prakhar Swarup, Prantik Sen, Debmalya Chakrabarty, K. V. Vijay Girish, Nikhil Bhave, Frederick Weber, Sambuddha Bhattacharya, Sri Garimella |
| 2025 | Speech-guided Grapheme-to-Phoneme Conversion for Cantonese Text-to-Speech. | Timothy Shin Heng Mak, King Yiu Suen, Albert Y. S. Lam |
| 2025 | SOMSRED-SVC: Sequential Output Modeling with Speaker Vector Constraints for Joint Multi-Talker Overlapped ASR and Speaker Diarization. | Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi, Ryo Masumura |
| 2025 | Unified Audio-Visual Modeling for Recognizing Which Face Spoke When and What in Multi-Talker Overlapped Speech and Video. | Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi, Ryo Masumura |
| 2025 | A Study on The Impact of Foundation Models on Automatic Depression Detection from Speech Signals. | Bubai Maji, Monorama Swain, Shazia Nasreen, Debabrata Majumdar, Rajlakshmi Guha, Aurobinda Routray, Anders Sgaard |
| 2025 | Chain-of-Thought Distillation with Fine-Grained Acoustic Cues for Speech Emotion Recognition. | Jialong Mai, Xiaofen Xing, Yangbiao Li, Xiangmin Xu |
| 2025 | AA-SLLM: An Acoustically Augmented Speech Large Language Model for Speech Emotion Recognition. | Jialong Mai, Xiaofen Xing, Weidong Chen, Yuanbo Fang, Xiangmin Xu |
| 2025 | CEREALES : a new dataset of Quebec French accented speech with applications to speech recognition. | Lucas Maison, Thomas Soulas, Marie-Jean Meurs |
| 2025 | Improving Linguistic Diversity of Large Language Models with Possibility Exploration Fine-Tuning. | Long Mai, Julie Carson-Berndsen |
| 2025 | Can Emotion Fool Anti-spoofing? | Aurosweta Mahapatra, Ismail Rasim Ulgen, Abinay Reddy Naini, Carlos Busso, Berrak Sisman |
| 2025 | Enabling the replicability of speech synthesis perceptual evaluations. | Sbastien Le Maguer, Gwnol Lecorv, Damien Lolive, Naomi Harte, Juraj Simko |
| 2025 | Multi-lingual and Zero-Shot Speech Recognition by Incorporating Classification of Language-Independent Articulatory Features. | Ryo Magoshi, Shinsuke Sakai, Jaeyoung Lee, Tatsuya Kawahara |
| 2025 | Joint Target-Speaker ASR and Activity Detection. | Chikara Maeda, Muhammad Shakeel, Yui Sudo |
| 2025 | Speaker Conditioning of Voice Activity Detection via Implicit Separation. | Matthew Maciejewski |
| 2025 | LLM-based phoneme-to-grapheme for phoneme-based speech recognition. | Te Ma, Min Bi, Saierdaer Yusuyin, Hao Huang, Zhijian Ou |
| 2025 | Assessment of L2 Oral Proficiency using Speech Large Language Models. | Rao Ma, Mengjie Qian, Siyuan Tang, Stefano Bann, Kate M. Knill, Mark J. F. Gales |
| 2025 | Temporal Modeling of Room Impulse Response Generation via Multi-Scale Autoregressive Learning. | Sheng Lyu, Yuemin Yu, Chenshu Wu |