| 2025 | Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning. | Yejin Jeon, Solee Im, Youngjae Kim, Gary Geunbae Lee |
| 2025 | Patient-Aware Feature Alignment for Robust Lung Sound Classification: Cohesion-Separation and Global Alignment Losses. | Seung Gyu Jeong, Seong Eun Kim |
| 2025 | NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data. | Tahir Javed, Kaushal Santosh Bhogale, Mitesh M. Khapra |
| 2025 | Prediction of listening effort ratings for habitual and clear-Lombard speech presented in noise. | Esther Janse, Chen Shen, Martin Cooke |
| 2025 | Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling. | Md Asif Jalal, Luca Remaggi, Vasileios Moschopoulos, Thanasis Kotsiopoulos, Vandana Rajan, Karthikeyan Saravanan, Anastasios Drosou, Junho Heo, Hyuk Oh, Seokyeong Jeong |
| 2025 | FaiST: A Benchmark Dataset for Fairness in Speech Technology. | Maliha Jahan, Yinglun Sun, Priyam Mazumdar, Zsuzsanna Fagyal, Thomas Thebaud, Jess Villalba, Mark Hasegawa-Johnson, Najim Dehak, Laureano Moro-Velzquez |
| 2025 | LombardTokenizer: Disentanglement and Control of Vocal Effort in a Neural Speech Codec. | Maxime Jacquelin, Mava Garnier, Laurent Girin, Rmy Vincent, Olivier Perrotin |
| 2025 | Selective Auditory Attention Decoding in Naturalistic Conversations Using EEG-Based Speech Envelope Tracking in Multi-Speaker Environments. | Gabriel Ivucic, Saurav Pahuja, Dashanka De Silva, Tanja Schultz |
| 2025 | Neural Speech Extraction with Human Feedback. | Malek Itani, Ashton Graves, Sefik Emre Eskimez, Shyamnath Gollakota |
| 2025 | Beyond Similarity Scoring: Detecting Entailment and Contradiction in Multilingual and Multimodal Contexts. | Othman Istaiteh, Salima Mdhaffar, Yannick Estve |
| 2025 | Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos. | Yuchi Ishikawa, Shota Nakada, Hokuto Munakata, Kazuhiro Saito, Tatsuya Komatsu, Yoshimitsu Aoki |
| 2025 | A Silent Speech Decoding System from EEG and EMG with Heterogenous Electrode Configurations. | Masakazu Inoue, Motoshige Sato, Kenichi Tomeoka, Nathania Nah, Eri Hatakeyama, Kai Arulkumaran, Ilya Horiguchi, Shuntaro Sasai |
| 2025 | PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs. | Sho Inoue, Shuai Wang, Haizhou Li |
| 2025 | Direction-Aware Neural Acoustic Fields for Few-Shot Interpolation of Ambisonic Impulse Responses. | Christopher Ick, Gordon Wichern, Yoshiki Masuyama, Franois G. Germain, Jonathan Le Roux |
| 2025 | Conformer-based Ultrasound-to-Speech Conversion. | Ibrahim Ibrahimov, Csaba Zaink, Gbor Gosztolya |
| 2025 | Voice-Based Dysphagia Detection: Leveraging Self-Supervised Speech Representation. | Injune Hwang, Jung-Min Kim, Ju Seok Ryu, Kyogu Lee |
| 2025 | Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition. | Cheng-Hung Hu, Yusuke Yasuda, Akifumi Yoshimoto, Tomoki Toda |
| 2025 | Speaker Normalization and Content Restoration for Zero-Shot Voice Conversion with Attention-Enhanced Discriminator. | Desheng Hu, Yang Xiang, Jian Lu, Xinhui Hu, Xinkang Xu |
| 2025 | On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition. | Shujie Hu, Xurong Xie, Mengzhe Geng, Jiajun Deng, Huimeng Wang, Guinan Li, Chengxi Deng, Tianzi Wang, Mingyu Cui, Helen Meng, Xunying Liu |
| 2025 | Does effortful speech production indicate communication difficulty caused by noise and hearing aid support? | Lena-Marie Huttner, Jeppe H. Christensen, Gitte Keidser, Tobias May, Torsten Dau, Sergi Rotger-Griful |
| 2025 | French schwa is not acoustically distinct from its two lexical neighbors // and /œ/. | Mathilde Hutin, Mlanie Lancien, Noam Faust |
| 2025 | HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement. | Amir Hussein, Sameer Khurana, Gordon Wichern, Franois G. Germain, Jonathan Le Roux |
| 2025 | Label Semantic-Driven Contrastive Learning for Speech Emotion Recognition. | Jiaxi Hu, Leyuan Qu, Haoxun Li, Taihao Li |
| 2025 | Word Level Timestamp Generation for Automatic Speech Recognition and Translation. | Ke Hu, Krishna C. Puvvada, Elena Rastorgueva, Zhehuai Chen, He Huang, Shuoyang Ding, Kunal Dhawan, Hainan Xu, Jagadeesh Balam, Boris Ginsburg |
| 2025 | Iterative Refinement, Not Training Objective, Makes HuBERT Behave Differently from wav2vec 2.0. | Robin Huo, Ewan Dunbar |