| 2025 | Exploratory Study of Filled Pauses in Ukrainian Language: Phonetic Properties of Filled Pauses. | Anna Havras, Carlos Mendes, Helena Moniz, Gueorgui Hristovsky, Joo Miranda |
| 2025 | Self-Supervised Models of Speech Processing for Haitian Creole. | William N. Havard, Renauld Govain, Benjamin Lecouteux, Emmanuel Schang |
| 2025 | DnR-nonverbal: Cinematic Audio Source Separation DatasetContaining Non-Verbal Sounds. | Takuya Hasumi, Yusuke Fujita |
| 2025 | Reddit FlairShare: A Human-Annotated Dataset of Gender-Progressive Online Discourse. | Carlos Hartmann |
| 2025 | SepVAC: Multitask Learning of Speaker Separation, Speaker Localization, Microphone Array Localization, and Room Acoustic Parameter Estimation in Various Acoustic Conditions. | Roland Hartanto, Sakriani Sakti, Koichi Shinoda |
| 2025 | Variability in performance across four generations of automatic speaker recognition systems. | Lauren Harrington, Vincent Hughes, Philip Harrison, Paul Foulkes, Jessica Wormald, Finnian Kelly, David van der Vloed |
| 2025 | Can ASR generate valid measures of child reading fluency? | Wieke Harmsen, Roeland van Hout, Catia Cucchiarini, Helmer Strik |
| 2025 | PAST: Phonetic-Acoustic Speech Tokenizer. | Nadav Har-Tuv, Or Tal, Yossi Adi |
| 2025 | L3C-DeepMFC: Low-Latency Low-Complexity Deep Marginal Feedback Cancellation with Closed-Loop Fine Tuning for Hearing Aids. | Fengyuan Hao, Brian C. J. Moore, Huiyong Zhang, Xiaodong Li, Chengshi Zheng |
| 2025 | WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation. | Lu Han, Junqi Zhao, Renhua Peng |
| 2025 | PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association. | Abdul Hannan, Muhammad Arslan Manzoor, Shah Nawaz, Muhammad Irzam Liaqat, Markus Schedl, Mubashir Noman |
| 2025 | An Effective Training Framework for Light-Weight Automatic Speech Recognition Models. | Abdul Hannan, Alessio Brutti, Shah Nawaz, Mubashir Noman |
| 2025 | Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization. | Jiangyu Han, Federico Landini, Johan Rohdin, Anna Silnova, Mireia Dez, Jan Cernock, Luks Burget |
| 2025 | Few-step Adversarial Schrdinger Bridge for Generative Speech Enhancement. | Seungu Han, Sungho Lee, Juheon Lee, Kyogu Lee |
| 2025 | CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-car Speech Separation with Distributed Heterogeneous Arrays. | Runduo Han, Yanxin Hu, Yihui Fu, Zihan Zhang, Yukai Jv, Li Chen, Lei Xie |
| 2025 | Automatic Speech Recognition for Low-Resourced Middle Eastern Languages. | Razhan Hameed, Sina Ahmadi, Hanah Hadi, Rico Sennrich |
| 2025 | Are loan sequences different from foreign sequences? A perception study with Japanese listeners on coronal obstruent - high front vowel sequences. | Silke Hamann, Andrea Alicehajic |
| 2025 | Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer. | Bence Mark Halpern, Thomas Tienkamp, Teja Rebernik, Rob J. J. H. van Son, Martijn Wieling, Defne Abur, Tomoki Toda |
| 2025 | Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition. | Jule Valendo Halim, Siyi Wang, Hong Jia, Ting Dang |
| 2025 | Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech. | Karl El Hajal, Enno Hermann, Sevada Hovsepyan, Mathew Magimai-Doss |
| 2025 | EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer. | Jiarui Hai, Yong Xu, Hao Zhang, Chenxing Li, Helin Wang, Mounya Elhilali, Dong Yu |
| 2025 | Phonetically-Augmented Discriminative Rescoring for Voice Search Error Correction. | Christophe Van Gysel, Maggie Wu, Lyan Verwimp, Caglar Tirkaz, Marco Bertola, Zhihong Lei, Youssef Oualil |
| 2025 | Deep learning based spatial aliasing reduction in beamforming for audio capture. | Mateusz Guzik, Giulio Cengarle, Daniel Arteaga |
| 2025 | Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN. | Yicheng Gu, Chaoren Wang, Zhizheng Wu, Lauri Juvela |
| 2025 | Audio-Based Classification and Geographic Regression of Austrian Dialects. | Lorenz Gutscher, Michael Pucher |