| 2025 | SpeechSEC: A Unified Multi-Task Framework for Speech Synthesis, Editing, and Continuation. | Liming Liang, Dongchao Yang, Xianwei Zhuang, Yuxin Xie, Luo Chen, Yuehan Jin, Yuexian Zou |
| 2025 | LightL2S: Ultra-Low Complexity Lip-to-Speech Synthesis for Multi-Speaker Scenarios. | Yifan Liang, Kang Yang, Fangkun Liu, Andong Li, Xiaodong Li, Chengshi Zheng |
| 2025 | FoleyMaster: High-Quality Video-to-Audio Synthesis via MLLM-Augmented Prompt Tuning and Joint Semantic-Temporal Adaptation. | Liming Liang, Luo Chen, Yuehan Jin, Xianwei Zhuang, Yuxin Xie, Yongkang Yin, Yuexian Zou |
| 2025 | Explainable Speech Emotion Recognition Through Attentive Pooling: Insights from Attention-Based Temporal Localization. | Tahitoa Leygue, Astrid Sabourin, Christian Bolzmacher, Sylvain Bouchigny, Margarita Anastassova, Quoc-Cuong Pham |
| 2025 | Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation. | Chenyang Le, Yinfeng Xia, Huiyan Li, Manhong Wang, Yutao Sun, Xingyang Ma, Yanmin Qian |
| 2025 | Towards the Objective Characterisation of Major Depressive Disorder Using Speech Data from a 12-week Observational Study with Daily Measurements. | Robert Lewis, Szymon Fedor, Nelson Hidalgo Julia, Joshua Curtiss, Jiyeon Kim, Noah Jones, David Mischoulon, Thomas F. Quatieri, Nicholas Cummins, Paola Pedrelli, Rosalind W. Picard |
| 2025 | Zero-Shot Mono-to-Binaural Speech Synthesis. | Alon Levkovitch, Julian Salazar, Soroosh Mariooryad, R. J. Skerry-Ryan, Nadav Bar, W. Bastiaan Kleijn, Eliya Nachmani |
| 2025 | An Exploration of Interpretable Deep Learning Models for the Assessment of Mild Cognitive Impairment. | Emma Cathrine Liisborg Leschly, Oliver Roesler, Michael Neumann, Jackson Liscombe, Abhishek Hosamath, Lakshmi Arbatti, Line H. Clemmensen, Melanie Ganz, Vikram Ramanarayanan |
| 2025 | Benchmarking Neural Speech Codec Intelligibility with SITool. | Anna Leschanowsky, Kishor Kayyar Lakshminarayana, Anjana Rajasekhar, Lyonel Behringer, Ibrahim Kilinc, Guillaume Fuchs, Emanul A. P. Habets |
| 2025 | Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction. | Thanathai Lertpetchpun, Tiantian Feng, Dani Byrd, Shrikanth Narayanan |
| 2025 | Leveraging Information Retrieval to Enhance Spoken Language Understanding Prompts in Few-Shot Learning. | Pierre Lepagnol, Sahar Ghannay, Thomas Gerald, Christophe Servan, Sophie Rosset |
| 2025 | SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification. | Tho Lepage, Rda Dehak |
| 2025 | Synthetic Data Generation for Phrase Break Prediction with Large Language Model. | Hoyeon Lee, Sejung Son, Ye-Eun Kang, Jong-Hwan Kim |
| 2025 | Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models. | Kyowoon Lee, Artyom Stitsyuk, Gunu Jho, Inchul Hwang, Jaesik Choi |
| 2025 | Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models. | Seung-jae Lee, Paul Hongsuck Seo |
| 2025 | Efficient Streaming TTS Acoustic Model with Depthwise RVQ Decoding Strategies in a Mamba Framework. | Joun Yeop Lee, Sangjun Park, Byoung Jin Choi, Ji-Hyun Lee, Min-Kyung Kim, Hoon-Young Cho |
| 2025 | Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation. | Jaejun Lee, Kyogu Lee |
| 2025 | DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization. | Geonyoung Lee, Geonhee Han, Paul Hongsuck Seo |
| 2025 | Articulatory Feature Prediction from Surface EMG during Speech Production. | Jihwan Lee, Kevin Huang, Kleanthis Avramidis, Simon Pistrosch, Monica Gonzlez Machorro, Yoonjeong Lee, Bjrn W. Schuller, Louis Goldstein, Shrikanth Narayanan |
| 2025 | Speech Enhancement based on cascaded two flows. | Seonggyu Lee, Sein Cheong, Sangwook Han, Kihyuk Kim, Jong Won Shin |
| 2025 | Multistage Universal Speech Enhancement System for URGENT Challenge. | Xiaohuai Le, Zhuangqi Chen, Siyu Sun, Xianjun Xia, Chuanzeng Huang |
| 2025 | Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods. | Laura Lechler, Chamran Moradi, Ivana Balic |
| 2025 | Real-Time Diffusion Buffer for Speech Enhancement On A Laptop. | Bunlong Lay, Rostilav Makarov, Timo Gerkmann |
| 2025 | Diffusion Buffer: Online Diffusion-based Speech Enhancement with Sub-Second Latency. | Bunlong Lay, Rostilav Makarov, Timo Gerkmann |
| 2025 | HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset. | Ryan Langman, Xuesong Yang, Paarth Neekhara, Shehzeen Hussain, Edresson Casanova, Evelina Bakhturina, Jason Li |