| 2025 | Speech power spectra: a window into neural oscillations in Parkinson's disease. | Sevada Hovsepyan, Mathew Magimai-Doss |
| 2025 | SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant. | Yixuan Hou, Heyang Liu, Yuhao Wang, Ziyang Cheng, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang |
| 2025 | Ranking and Selection of Bias Words for Contextual Bias Speech Recognition. | Haoxiang Hou, Xun Gong, Wangyou Zhang, Wei Wang, Yanmin Qian |
| 2025 | Why is children's ASR so difficult? Analyzing children's phonological error patterns using SSL-based phoneme recognizers. | Koharu Horii, Naohiro Tawara, Atsunori Ogawa, Shoko Araki |
| 2025 | Pretraining Multi-Speaker Identification for Neural Speaker Diarization. | Shota Horiguchi, Atsushi Ando, Naohiro Tawara, Marc Delcroix |
| 2025 | Mitigating Non-Target Speaker Bias in Guided Speaker Embedding. | Shota Horiguchi, Takanori Ashihara, Marc Delcroix, Atsushi Ando, Naohiro Tawara |
| 2025 | FUSE-MOS: Fusion of Speech Embeddings for MOS Prediction with Uncertainty Quantification. | Enjamamul Hoq, Nikhil Gupta, Danielle Omondi, Ifeoma Nwogu |
| 2025 | FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents. | Satu Hopponen, Tomi Kinnunen, Alexandre Nikolaev, Rosa Gonzlez Hautamki, Lauri Tavi, Einar Meister |
| 2025 | Voices of 'cyborg awesomeness': Posthuman embodiment of nonbinary gender expression in AI speech technologies. | Maxwell Hope, va Szkely |
| 2025 | Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization. | Luong Ho, Khanh Le, Vinh Pham, Bao Nguyen, Tan Tran, Duc Chau |
| 2025 | Using and comprehending language in face-to-face conversation. | Judith Holler |
| 2025 | Revisiting WFST-based Hybrid Japanese Speech Recognition System for Individuals with Organic Speech Disorders. | Naoki Hojo, Ryoichi Takashima, Chihiro Sugiyama, Nobukazu Tanaka, Kanji Nohara, Kazunori Nozaki, Tetsuya Takiguchi |
| 2025 | Hearing deficits of transformer-based ASR for anechoic and spatial signals. | Dirk Eike Hoffner, Simon Weihe, Thomas Brand, Bernd T. Meyer |
| 2025 | Enhancing Transcripts of Open-Source Automatic Speech Recognition Models Through Fine-Tuning with Laughter and Speech-Laugh. | Phuoc Hoang Ho, Dragos Alexandru Balan, Dirk K. J. Heylen, Khiet P. Truong |
| 2025 | Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment. | Long-Vu Hoang, Tuan Nguyen, Huy Dat Tran |
| 2025 | Hybrid Data Sampling for ASR: Integrating Acoustic Diversity and Transcription Uncertainty. | Komei Hiruta, Yosuke Yamano, Hideaki Tamori |
| 2025 | SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition. | Yuta Hirano, Sakriani Sakti |
| 2025 | Analyzing the Importance of Blank for CTC-Based Knowledge Distillation. | Benedikt Hilmes, Nick Rossenbach, Ralf Schlter |
| 2025 | End-to-End Speech Translation Guided by Robust Translation Capability of Large Language Model. | Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi |
| 2025 | TTMBA: Towards Text To Multiple Sources Binaural Audio Generation. | Yuxuan He, Xiaoran Yang, Ningning Pan, Gongping Huang |
| 2025 | CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models. | Jiajun He, Naoki Sawada, Koichi Miyazaki, Tomoki Toda |
| 2025 | Acoustic similarities, articulatory uniqueness: Speech production mechanisms in individuals with congenital lip paralysis. | Anne Hermes, Ivana Didirkov, Philipp Buech, Gilles Vannuscorps |
| 2025 | Gaze-Enhanced Multimodal Turn-Taking Prediction in Triadic Conversations. | Seongsil Heo, Christi Miller, Calvin Murdock, Michael J. Proulx |
| 2025 | GIA-MIC: Multimodal Emotion Recognition with Gated Interactive Attention and Modality-Invariant Learning Constraints. | Jiajun He, Jinyi Mi, Tomoki Toda |
| 2025 | Factors affecting the in-context learning abilities of LLMs for dialogue state tracking. | Pradyoth Hegde, Santosh Kesiraju, Jan Svec, Simon Sedlcek, Bolaji Yusuf, Oldrich Plchot, Deepak K. T, Jan Cernock |