| 2025 | DLF-EEND: Dynamic Layer Fusion for End-to-End Speaker Diarization. | Wooil Kim, Bongsu Jung |
| 2025 | Data Augmentation using Speech Synthesis for Speaker-Independent Dysarthria Severity Classification. | Minseop Kim, Minsu Han, Seokyoung Hong, Myoung-wan Koo |
| 2025 | Fully End-to-end Streaming Open-vocabulary Keyword Spotting with W-CTC Forced Alignment. | Dohyun Kim, Jiwook Hwang |
| 2025 | A Hybrid Approach to Combining Role Diarization with ASR for Professional Conversations. | Bongjun Kim, Arindam Ghosh, Mark C. Fuhs, Anurag Chowdhury, Deblin Bagchi, Monika Woszczyna |
| 2025 | Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech. | Nam-Gyu Kim, Deok-Hyeon Cho, Seung-Bin Kim, Seong-Whan Lee |
| 2025 | Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework. | Yunsik Kim, Yoonyoung Chung |
| 2025 | Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis. | Minsu Kim, Pingchuan Ma, Honglie Chen, Stavros Petridis, Maja Pantic |
| 2025 | Steering Deep Non-Linear Spatially Selective Filters for Weakly Guided Extraction of Moving Speakers in Dynamic Scenarios. | Jakob Kienegger, Timo Gerkmann |
| 2025 | AttentiveMOS: A Lightweight Attention-Only Model forSpeech Quality Prediction. | Imran E. Kibria, Donald S. Williamson |
| 2025 | Factorized RVQ-GAN For Disentangled Speech Tokenization. | Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Bobos, Juraj Novosad, Peter Gazdik, Ellen Zhang, Zili Huang, Amir Hussein, Ricard Marxer, Yoshiki Masuyama, Ryo Aihara, Chiori Hori, Franois G. Germain, Gordon Wichern, Jonathan Le Roux |
| 2025 | BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention. | Yassine El Kheir, Tim Polzehl, Sebastian Mller |
| 2025 | Towards a Unified Benchmark for Arabic Pronunciation Assessment: Qur'anic Recitation as Case Study. | Yassine El Kheir, Omnia Ibrahim, Amit Meghanani, Nada Almarwani, Hawau Olamide Toyin, Sadeen Alharbi, Modar Alfadly, Lamya Alkanhal, Ibrahim Selim, Shehab Elbatal, Salima Mdhaffar, Thomas Hain, Yasser Hifny, Mostafa Shahin, Ahmed Ali |
| 2025 | Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings. | Owais Mujtaba Khanday, Pablo Rodrguez San Esteban, Zubair Ahmad Lone, Marc Ouellet, Jos Andrs Gonzlez Lpez |
| 2025 | Optimizing Pause Context in Fine-Tuning Pre-trained Large Language Models for Dementia Detection. | Xiaoquan Ke, Man-Wai Mak, Helen Meng |
| 2025 | Improving User Impression of Spoken Dialogue Systems by Controlling Para-linguistic Expression Based on Intimacy. | Shoki Kawanishi, Akinori Ito, Yuya Chiba, Takashi Nose |
| 2025 | BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing. | Masaya Kawamura, Takuya Hasumi, Yuma Shirahata, Ryuichi Yamamoto |
| 2025 | From Scarcity to Sufficiency: Speech Recognition Pipeline for Zero-resource Language. | Nikolay Karpov, Sofia Kostandian, Nune Tadevosyan, Alexan Ayrapetyan, Andrei Andrusenko, Ara Yeroyan, Mher Yerznkanyan, Vitaly Lavrukhin |
| 2025 | Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy. | Elvir Karimov, Alexander Varlamov, Danil Ivanov, Dmitrii Korzh, Oleg Rogov |
| 2025 | Pick and Summarize: Integrating Extractive and Abstractive Speech Summarization. | Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Ryo Fukuda, William Chen, Shinji Watanabe |
| 2025 | When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds. | Minsu Kang, Seolhee Lee, Choonghyeon Lee, Namhyun Cho |
| 2025 | Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech. | Wonjune Kang, Junteng Jia, Chunyang Wu, Wei Zhou, Egor Lakomkin, Yashesh Gaur, Leda Sari, Suyoun Kim, Ke Li, Jay Mahadeokar, Ozlem Kalinli |
| 2025 | Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation. | Fang Kang, Yin Cao, Haoyu Chen |
| 2025 | Vocoder-Projected Feature Discriminator. | Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo |
| 2025 | FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation. | Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo |
| 2025 | Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models. | Shunsuke Kando, Yusuke Miyao, Shinnosuke Takamichi |