| 2025 | DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration. | Sanberk Serbest, Tijana Stojkovic, Milos Cernak, Andrew Harper |
| 2025 | FaVC: A Validated, Transcribed, Parallel Farsi Speech Dataset for Voice Conversion. | Mina Serajian, Saeed Najafzadeh Rahaghi, Hadi Veisi, Saman Haratizadeh |
| 2025 | Enhancing Target-speaker Automatic Speech Recognition Using Multiple Speaker Embedding Extractors with Virtual Speaker Embedding. | Ju-Seok Seong, Jeong-Hwan Choi, Ye-Rin Jeoung, Ilseok Kim, Joon-Hyuk Chang |
| 2025 | Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs. | Simon Sedlcek, Bolaji Yusuf, Jan Svec, Pradyoth Hegde, Santosh Kesiraju, Oldrich Plchot, Jan Cernock |
| 2025 | PredTrAD - Prediction-based Transformer for Anomaly Detection in Multivariate Time Series Data. | Jan Schuster, Alexander Wlfel, Fabian Brunner, Christian Bergler |
| 2025 | Processing of grammatical information in cochlear implant simulated speech by German adult listeners. | Atty Schouwenaars, Esther Ruigendijk |
| 2025 | DiffMV-ETS: Diffusion-based Multi-Voice Electromyography-to-Speech Conversion using Speaker-Independent Speech Training Targets. | Kevin Scheck, Tom Dombeck, Zhao Ren, Peter Wu, Michael Wand, Tanja Schultz |
| 2025 | Hear Me Out: Interactive evaluation and bias discovery platform for speech-to-speech conversational AI. | Shree Harsha Bokkahalli Satish, Gustav Eje Henter, va Szkely |
| 2025 | Pitch Accent Detection improves Pretrained Automatic Speech Recognition. | David Sasu, Natalie Schluter |
| 2025 | Enhancing Speech Instruction Understanding and Disambiguation in Robotics via Speech Prosody. | David Sasu, Benedict Quartey, Kweku Andoh Yamoah, Natalie Schluter |
| 2025 | LHCP-ASR: An English Speech Corpus of High-Energy Particle Physics Talks for Narrow-Domain ASR Benchmarking. | Jaume Santamaria-Jorda, Pablo Segovia-Martnez, Gonal V. Garcs Daz-Muno, Joan Albert Silvestre-Cerd, Adri Gimnez, Rubn Gaspar Aparicio, Ren Fernndez Snchez, Jorge Civera, Albert Sanchs, Alfons Juan |
| 2025 | Rasmalai : Resources for Adaptive Speech Modeling in IndiAn Languages with Accents and Intonations. | Ashwin Sankar, Yoach Lacombe, Sherry Thomas, Praveen Srinivasa Varadhan, Sanchit Gandhi, Mitesh M. Khapra |
| 2025 | Adversarial Attacks on Text-dependent Speaker Verification System. | Sreekanth Sankala, Venkatesh Parvathala, Ramesh Gundluru, K. Sri Rama Murty |
| 2025 | Physiologically-Informed Feature Analysis of Acquired Speech Disorders for Stroke Assessment. | Giulia Sanguedolce, Jn Gunason, Dragos-Cristian Gruia, Emilie D'Olne, Fatemeh Geranmayeh, Patrick A. Naylor |
| 2025 | Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information. | Nicholas Sanders, Yuanchao Li, Korin Richmond, Simon King |
| 2025 | Can We Reconstruct a Dysarthric Voice with the Large Speech Model Parler TTS? | Ariadna Sanchez, Simon King |
| 2025 | A Cookbook for Community-driven Data Collection of Impaired Speech in Low-Resource Languages. | Sumaya Ahmed Salihs, Isaac Wiafe, Jamal-Deen Abdulai, Elikem Doe Atsakpo, Gifty Ayoka, Richard Cave, Akon Obu Ekpezu, Catherine Holloway, Katrin Tomanek, Fiifi Baffoe Payin Winful |
| 2025 | Multichannel Keyword Spotting for Noisy Conditions. | Dzmitry Saladukha, Ivan Koriabkin, Kanstantsin Artsiom, Aliaksei Rak, Nikita Ryzhikov |
| 2025 | Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition. | Asahi Sakuma, Hiroaki Sato, Ryuga Sugano, Tadashi Kumano, Yoshihiko Kawai, Tetsuji Ogawa |
| 2025 | Interspeech 2025 URGENT Speech Enhancement Challenge. | Kohei Saijo, Wangyou Zhang, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar, Marvin Sach, Yihui Fu, Wei Wang, Tim Fingscheidt, Shinji Watanabe |
| 2025 | Visual features of the oral region in Polish sibilants produced by children with various sibilance patterns. | Agata Sage, Zuzanna Miodonska, Michal Krecichwost, Ewa Kwasniok, Pawel Badura |
| 2025 | Bringing Interpretability to Neural Audio Codecs. | Samir Sadok, Julien Hauret, ric Bavu |
| 2025 | Unified Microphone Conversion: Many-to-Many Device Mapping via Feature-wise Linear Modulation. | Myeonghoon Ryu, Hongseok Oh, Suji Lee, Han Park |
| 2025 | Pitch Contour Model (PCM) with Transformer Cross-Attention for Speech Emotion Recognition. | Minji Ryu, Jihyeon Hur, Sung Heuk Kim, Gahgene Gweon |
| 2025 | Dhvani: A Weakly-supervised Phonemic Error Detection and Personalized Feedback System for Hindi. | Arnav Rustagi, Satvik Bajpai, Nimrat Kaur, Siddharth |