| 2025 | J-SPAW: Japanese speaker verification and spoofing attacks recorded in-the-wild dataset. | Sayaka Shiota, Suzuka Horie, Kouta Kanno, Shinnosuke Takamichi |
| 2025 | Generating Consistent Prosodic Patterns from Open-Source TTS Systems. | Ha Eun Shim, Olivia Yung, Paige Tutts, Boey Kwan, Angelica Lim, Yue Wang, H. Henny Yeung |
| 2025 | Advancing Emotion Recognition via Ensemble Learning: Integrating Speech, Context, and Text Representations. | Xiaohan Shi, Jinyi Mi, Xingfeng Li, Tomoki Toda |
| 2025 | Universal Preference-Score-based Pairwise Speech Quality Assessment. | Yufei Shi, Yang Ai, Zhen-Hua Ling |
| 2025 | Speaker-Aware Multi-Task Learning for Speech Emotion Recognition. | Xiaohan Shi, Xingfeng Li, Tomoki Toda |
| 2025 | Who, When, and What: Leveraging the "Three Ws" Concept for Emotion Recognition in Conversation. | Xiaohan Shi, Xingfeng Li, Tomoki Toda |
| 2025 | Tungna In Live Performance: An Implementation Of Interactive Artistic Text-To-Voice. | Victor Shepardson, Jonathan Reus, Thor Magnusson |
| 2025 | ARiSE: Auto-Regressive Multi-Channel Speech Enhancement. | Pengjie Shen, Xueliang Zhang, Zhong-Qiu Wang |
| 2025 | On the reliability of feature attribution methods for speech classification. | Gaofei Shen, Hosein Mohebbi, Arianna Bisazza, Afra Alishahi, Grzegorz Chrupala |
| 2025 | Scalable Spontaneous Speech Dataset (SSSD): Crowdsourcing Data Collection to Promote Dialogue Research. | Zaid Sheikh, Shuichiro Shimizu, Siddhant Arora, Jiatong Shi, Samuele Cornell, Xinjian Li, Shinji Watanabe |
| 2025 | Towards Secure User Authentication for Headphones via In-Ear or In-Earcup Microphones. | N. Shashaank, Xiao Quan, Andrew Kaluzny, Leonard Varghese, Marko Stamenovic, Chuan-Che Huang |
| 2025 | Enhancing Syllabic Recognition via Speech-EEG Phase Analysis and Non-Activity State Modeling. | Rini A. Sharon, Hema A. Murthy |
| 2025 | Boosting StoRM Convergence with Metric Guidance and Non-uniform State-Sampling for Optimal Dereverberation. | Chandra Mohan Sharma, Arnab Kumar Roy, Anupam Mandal, Prasanta Kumar Ghosh, Prasanna Kumar Kr |
| 2025 | A real-time MRI study on asymmetry in velum dynamics during VCV production with nasal sounds. | Chetan Sharma, Vaishnavi Chandwanshi, Shreya Shrikant Karkun, Aditya Anand Gupta, Prasanta Kumar Ghosh |
| 2025 | Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR. | Mingchen Shao, Xinfa Zhu, Chengyou Wang, Bingshen Mu, Hai Li, Ying Yan, Junhui Liu, Danming Xie, Lei Xie |
| 2025 | Lexical stress affects lenition: The case of Italian palato-alveolar affricates. | Bowei Shao, Philipp Buech, Anne Hermes, Maria Giavazzi |
| 2025 | CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR. | Natarajan Balaji Shankar, Zilai Wang, Kaiyuan Zhang, Mohan Shi, Abeer Alwan |
| 2025 | Neuro2Semantic: A Transfer Learning Framework for Semantic Reconstruction of Continuous Language from Human Intracranial EEG. | Siavash Shams, Richard J. Antonello, Gavin Mischler, Stephan Bickel, Ashesh D. Mehta, Nima Mesgarani |
| 2025 | Harnessing Text-to-Speech Voice Cloning Models for Improved Audiological Speech Assessment. | Lidea Shahidi, Erdem Baha Topbas, Thu Ngan Dang, Tobias Goehring |
| 2025 | NAM-to-Speech Conversion with Multitask-Enhanced Autoregressive Models. | Neil Shah, Shirish Karande, Vineet Gandhi |
| 2025 | Spoken Question Answering for Visual Queries. | Nimrod Shabtay, Zvi Kons, Avihu Dekel, Hagai Aronowitz, Ron Hoory, Assaf Arbelle |
| 2025 | Assessment of the synthetic quality and controllability of laughing onset in speech-laugh synthesis. | Ryo Setoguchi, Yoshiko Arimoto |
| 2025 | MTSE: Multi-Target Speaker Extraction for Conversation Scenarios. | Thomas Serre, Mathieu Fontaine, Eric Benhaim, Slim Essid |
| 2025 | CommissionsQC: a Qubec French Speech Corpus for Automatic Speech Recognition. | Coralie Serrand, Amira Morsli, Gilles Boulianne |
| 2025 | Automatic Speech Recognition Biases in Newcastle English: an Error Analysis. | Dana Serditova, Kevin Tang, Jochen Steffens |