| 2025 | Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages. | Rishabh Ranjan, Ayinala Likhith, Mayank Vatsa, Richa Singh |
| 2025 | Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering. | Pradeep Rangappa, Andrs Carofilis, Jeena J. Prakash, Shashi Kumar, Sergio Burdisso, Srikanth R. Madikeri, Esa Villatoro-Tello, Bidisha Sharma, Petr Motlcek, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke |
| 2025 | Bilingual Speakers Exhibit Cognitive Fatigue: A Speech Disfluencies Case Study on Research Talks. | Ashwin Ram, Marisol Muoz, Zoi Gkalitsiou, Alexandros G. Dimakis |
| 2025 | Oral Reading Errors by Grade 3 Children in Indian Schools: A Hindi-English Perspective. | Sneha Raman, Preeti Rao |
| 2025 | End-to-End Indian Language Dubbing with Zero-Shot Speaker Preservation. | Giri Raju, Sandeep Konam |
| 2025 | ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems. | Anand Kumar Rai, Satyam Rahangdale, Utkarsh Anand, Animesh Mukherjee |
| 2025 | Characterization of voice cue sensitivity and vocal emotion recognition across the adult lifespan. | Laura Rachman, Deniz Baskent |
| 2025 | VisualSpeech: Enhancing Prosody Modeling in TTS Using Video. | Shumin Que, Anton Ragni |
| 2025 | Exploring Generative Error Correction for Dysarthric Speech Recognition. | Moreno La Quatra, Alkis Koudounas, Valerio Mario Salerno, Sabato Marco Siniscalchi |
| 2025 | PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts. | Tianhua Qi, Shiyan Wang, Cheng Lu, Tengfei Song, Hao Yang, Zhanglin Wu, Wenming Zheng |
| 2025 | Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming. | Chengyuan Qin, Wenmeng Xiong, Jing Zhou, Maoshen Jia, Changchun Bao |
| 2025 | Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction. | Mengjie Qian, Rao Ma, Stefano Bann, Kate M. Knill, Mark J. F. Gales |
| 2025 | Representation of Perceived Prosodic Similarity of Conversational Feedback. | Livia Qian, Carol Figueroa, Gabriel Skantze |
| 2025 | Empowering Large Language Models for End-to-End Speech Translation Leveraging Synthetic Data. | Yu Pu, Xiaoqian Liu, Guangyu Zhang, Zheng Yan, Wei-Qiang Zhang, Xie Chen |
| 2025 | An approach to measuring the performance of Automatic Speech Recognition(ASR) models in the context of Large Language Model(LLM) powered applications. | Sujith Pulikodan, Sahapthan K, Prasanta Kumar Ghosh, Visruth Sanka, Nihar Desai |
| 2025 | Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM. | Dariia Puhach, Amir H. Payberah, va Szkely |
| 2025 | Rhotic Articulation in Australian English: Insights from MRI. | Michael Proctor, Tnde Szalay, Tharinda Piyadasa, Craig T. Jin, Naeim Sanaei, Amelia Gully, David Waddington, Sheryl Foster, Kirrie J. Ballard |
| 2025 | Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation. | Gowtham Premananth, Philip Resnik, Sonia Bansal, Deanna L. Kelly, Carol Y. Espy-Wilson |
| 2025 | Analyzing the Impact of Accent on English Speech: Acoustic and Articulatory Perspectives. | Gowtham Premananth, Vinith Kugathasan, Carol Y. Espy-Wilson |
| 2025 | Explainable Depression Detection using Masked Hard Instance Mining. | Patawee Prakrankamanant, Shinji Watanabe, Ekapol Chuangsuwanich |
| 2025 | Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM. | Jeena J. Prakash, Blessingh Kumar, Kadri Hacioglu, Bidisha Sharma, Sindhuja Gopalan, Malolan Chetlur, Shankar Venkatesan, Andreas Stolcke |
| 2025 | End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data. | Aishwarya Pothula, Bhavana Akkiraju, Srihari Bandarupalli, Charan Devarkonda, Santosh Kesiraju, Anil Kumar Vuppala |
| 2025 | Evaluating the Effectiveness of Pre-Trained Audio Embeddings for Classification of Parkinson's Disease Speech Data. | Emmy Postma, Cristian Tejedor Garca |
| 2025 | Learning Optimal Prosody Embedding Codebook based on F0 and Energy. | David Portes, Ales Hork |
| 2025 | Tracking /r/ Deletion: Forced Alignment of Pronunciation Variants and Sociophonetic Insights into Post-Obstruent Final /r/ in French. | Anisia Popescu, Lori Lamel, Marc Evrard, Ioana Vasilescu |