| 2025 | Intelligibility Prediction for Time-Modified Speech Signals Using Spectro-Temporal Modulation Features. | Aymen Bashir, Haolan Wang, Amin Edraki, Wai-Yip Chan, Jesper Jensen |
| 2025 | PhonemeFake: Redefining Deepfake Realism with Language-Driven Segmental Manipulation and Adaptive Bilevel Detection. | Oguzhan Baser, Ahmet Ege Tanriverdi, Sriram Vishwanath, Sandeep Chinchali |
| 2025 | WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing. | Oguzhan Baser, Ahmet Ege Tanriverdi, Kaan Kale, Sandeep Chinchali, Sriram Vishwanath |
| 2025 | Analysis of ABC Frontend Audio Systems for the NIST-SRE24. | Sara Barahona, Anna Silnova, Ladislav Mosner, Junyi Peng, Oldrich Plchot, Johan Rohdin, Lin Zhang, Jiangyu Han, Petr Plka, Federico Landini, Luks Burget, Themos Stafylakis, Sandro Cumani, Dominik Bobos, Miroslav Hlavcek, Martin Kodovsky, Toms Pavlcek |
| 2025 | Frequency-Domain Enhanced Extreme Bandwidth Extension Network with ICCRN for Superior Speech Quality. | Hongtao Bao, Xueliang Zhang |
| 2025 | Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings. | Jesuraj Bandekar, Prasanta Kumar Ghosh |
| 2025 | SMARTMOS: Modeling Subjective Audio Quality Evaluation for Real-Time Applications. | Sivakumar Balasubramanian, Jose Antonio Jimenez Amador, Kaustubh Kalgaonkar, King-Wei Hor, Sriram Srinivasan |
| 2025 | Influence of Proficiency and L2 Experience on Dynamic Spectral Cue Utilization in L2 Vowel Perception and Production. | Linda Bakkouche, Brechtje Post |
| 2025 | Finding the Human Voice in AI: Insights on the Perception of AI-Voice Clones from Naturalness and Similarity Ratings. | Linda Bakkouche, Charles McGhee, Emily Lau, Stephanie Cooper, Xinbing Luo, Madeleine Rees, Kai Alter, Brechtje Post, Julia Schwarz |
| 2025 | Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data. | Qibing Bai, Sho Inoue, Shuai Wang, Zhongjie Jiang, Yannan Wang, Haizhou Li |
| 2025 | Co-Speech Motion for Virtual Agents in Dialogue Using LLM-Driven Primitive Action Selection. | Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino |
| 2025 | Rapport-Building Dialogue Strategies for Deeper Connection: Integrating Proactive Behavior, Personalization, and Aizuchi Backchannels. | Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino |
| 2025 | Mixture of LoRA Experts for Low-Resourced Multi-Accent Automatic Speech Recognition. | Raphal Bagat, Irina Illina, Emmanuel Vincent |
| 2025 | LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech. | Niyati Bafna, Matthew Wiesner |
| 2025 | Reconstruction of the Complete Vocal Tract Contour Through Acoustic to Articulatory Inversion Using Real-Time MRI Data. | Sofiane Azzouz, Pierre-Andr Vuissoz, Yves Laprie |
| 2025 | Deep-Simplex Multichannel Speech Separation. | Tzlil Avidan, Bracha Laufer-Goldshtein |
| 2025 | From Weak Labels to Strong Results: Utilizing 5, 000 Hours of Noisy Classroom Transcripts with Minimal Accurate Data. | Ahmed Adel Attia, Dorottya Demszky, Jing Liu, Carol Y. Espy-Wilson |
| 2025 | Analysis of Semantic and Acoustic Token Variability Across Speech, Music, and Audio Domains. | Takanori Ashihara, Marc Delcroix, Tsubasa Ochiai, Kohei Matsuura, Shota Horiguchi |
| 2025 | ATMM-SAGA: Alternating Training for Multi-Module with Score-Aware Gated Attention SASV system. | Amro Asali, Yehuda Ben-Shimol, Itshak Lapidot |
| 2025 | Chain-of-Thought Training for Open E2E Spoken Dialogue Systems. | Siddhant Arora, Jinchuan Tian, Hayato Futami, Jee-weon Jung, Jiatong Shi, Yosuke Kashiwagi, Emiru Tsunoo, Shinji Watanabe |
| 2025 | Evaluating Large Language Models in Data Generation for Low-Resource Scenarios: A Case Study on Question Answering. | Ebru Arisoy, Merve nl Menevse, Yusufcan Manav, Arzucan zgr |
| 2025 | Coping with segmental-prosodic incongruity in spoken word recognition in Japanese. | Terumichi Ariga |
| 2025 | Robot-assisted Recognition of Vocal Emotions in Pseudospeech for Cochlear Implanted Adolescents. | Gloria Araiza-Illan, Luke Meyer, Bert Maat, Deniz Baskent |
| 2025 | Vocal-tract model with two directions: Static design for a dummy head and dynamic design for a speaking machine. | Takayuki Arai |
| 2025 | A Naturally Elicited Multimodal Stress Database and Speech Breathing Based Stress Detection. | Karumannil Mohamed Ismail Yasar Arafath, Mohammed Abeer K. C., Aurobinda Routray |