| 2025 | VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech. | Harm Lameris, Joakim Gustafson, va Szkely |
| 2025 | Unified Variational and Physics-aware Model for Room Impulse Response Estimation. | Louis Lalay, Mathieu Fontaine, Roland Badeau |
| 2025 | Investigating Affect Mining Techniques for Annotation Sample Selection in the Creation of Finnish Affective Speech Corpus. | Kalle Lahtinen, Einari Vaaras, Liisa Mustanoja, Okko Rsnen |
| 2025 | Parameter-Efficient Fine-Tuning for Low-Resource Text-to-Speech via Cross-Lingual Continual Learning. | Ki-Joong Kwon, Jun-Ho So, Sang-Hoon Lee |
| 2025 | Speaker-specific Patterns of Phonetic Covariation in Korean Word-medial Stops and the Role of Phonological and Morphological Contexts. | Chloe D. Kwon |
| 2025 | GigaAM: Efficient Self-Supervised Learner for Speech Recognition. | Aleksandr Kutsakov, Alexandr Maximenko, Georgii Gospodinov, Pavel Bogomolov, Fyodor Minkin |
| 2025 | Automatic Dialectal Transcription: An Evaluation on Finnish and Norwegian. | Olli Kuparinen |
| 2025 | Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier. | Tarek Kunze, Marianne Mtais, Hadrien Titeux, Lucas Elbert, Joseph Coffey, Emmanuel Dupoux, Alejandrina Cristi, Marvin Lavechin |
| 2025 | SGED-Probe: Probing E2E ASR decoder and aligner for spoken grammar error detection under three speaking practice conditions. | Chowdam Venkata Thirumala Kumar, Chiranjeevi Yarra |
| 2025 | DRI-GAN: A Novel Dual Real Input GAN with Triplet Loss for Cross-Lingual and Noisy SLU. | Ankit Kumar, Munir Georges |
| 2025 | Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion. | Saurabh Kumar, Amartyaveer, Prasanta Kumar Ghosh |
| 2025 | ArticulateX: End-to-End Monolingual Speech Translation in Articulator Space. | Vishal Kumar, Vinayak Abrol |
| 2025 | Children's Voice Privacy: First Steps and Emerging Challenges. | Ajinkya Kulkarni, Francisco Teixeira, Enno Hermann, Thomas Rolland, Isabel Trancoso, Mathew Magimai-Doss |
| 2025 | Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion. | Ajinkya Kulkarni, Sandipana Dowerah, Tanel Alume, Mathew Magimai-Doss |
| 2025 | xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement. | Nikolai Lund Khne, Jan stergaard, Jesper Jensen, Zheng-Hua Tan |
| 2025 | Towards Frame-level Quality Predictions of Synthetic Speech. | Michael Kuhlmann, Fritz Seebauer, Petra Wagner, Reinhold Haeb-Umbach |
| 2025 | Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples. | Chun-Yi Kuan, Hung-yi Lee |
| 2025 | Leveraging Text and Speech Processing for Suicide Risk Classification in Chinese Adolescents. | Justyna Krzywdziak, Bartlomiej Eljasiak, Joanna Stepien, Michal Swiatek, Agnieszka Pruszek |
| 2025 | Extending the Fongbe to French Speech Translation Corpus: resources, models and benchmark. | D. Fortune Kponou, Salima Mdhaffar, Frjus A. A. Laleye, Eugne C. Ezin, Yannick Estve |
| 2025 | Synthetic Speech Source Tracing using Metric Learning. | Dimitrios Koutsianos, Stavros Zacharopoulos, Yannis Panagakis, Themos Stafylakis |
| 2025 | "Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding. | Alkis Koudounas, Claudio Savelli, Flavio Giobergia, Elena Baralis |
| 2025 | "KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding. | Alkis Koudounas, Moreno La Quatra, Eliana Pastor, Sabato Marco Siniscalchi, Elena Baralis |
| 2025 | MVP: Multi-source Voice Pathology detection. | Alkis Koudounas, Moreno La Quatra, Gabriele Ciravegna, Marco Fantini, Erika Crosetti, Giovanni Succo, Tania Cerquitelli, Sabato Marco Siniscalchi, Elena Baralis |
| 2025 | Multimodal Speech-Based Biomarkers Outperform the ALS Functional Rating Scale in Predicting Individual Disease Progression in ALS. | Hardik Kothare, Michael Neumann, Vikram Ramanarayanan |
| 2025 | Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform. | Xiangzhu Kong, Hao Huang, Zhijian Ou |