| 2025 | Visual Cues Support Robust Turn-taking Prediction in Noise. | Sam O'Connor Russell, Naomi Harte |
| 2025 | Intelligibility of Text-to-Speech Systems for Mathematical Expressions. | Sujoy Roychowdhury, Ranjani Hosakere Gireesha, Sumit Soman, Nishtha Paul, Subhadip Bandyopadhyay, Siddhanth Iyengar |
| 2025 | Structured pruning for efficient systolic array accelerated cascade Speech-to-Text Translation. | Jean-Luc Rouas, Charles Brazier, Leila Ben Letaifa, Rafael Medina, Pedro Palacios, David Atienza, Giovanni Ansaloni |
| 2025 | Conveying Gender Through Speech: Insights from Trans Men. | Alice Ross, Cliodhna Hughes, Eddie L. Ungless, Catherine Lai |
| 2025 | Running Conventional Automatic Speech Recognition on Memristor Hardware: A Simulated Approach. | Nick Rossenbach, Benedikt Hilmes, Leon Brackmann, Moritz Gunz, Ralf Schlter |
| 2025 | Advancing Pediatric ASR: The Role of Voice Generation in Disordered Speech. | Karen Rosero, Ali N. Salman, Shreeram Suresh Chandra, Berrak Sisman, Cortney Van't Slot, Alex A. Kane, Rami R. Hallac, Carlos Busso |
| 2025 | In-context learning capabilities of Large Language Models to detect suicide risk among adolescents from speech transcripts. | Filomene Roquefort, Alexandre Ducorroy, Rachid Riad |
| 2025 | TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement Network. | Xiaobin Rong, Dahan Wang, Qinwen Hu, Yushi Wang, Yuxiang Hu, Jing Lu |
| 2025 | Pre-aspiration in Iceland Is Conditioned by Gender/Sex. | Meike Rommel, Msa Hejn, Nicole Deh |
| 2025 | SCRIBAL: A Digital Transcription Tool in Higher Education. | Javier Romn, Pol Pastells, Mauro Vzquez Chas, Clara Puigvents, Montserrat Nofre, Mariona Taul, Mireia Farrs |
| 2025 | Exploring Shared-Weight Mechanisms in Transformer and Conformer Architectures for Automatic Speech Recognition. | Thomas Rolland, Alberto Abad |
| 2025 | Improving Bird Classification with Primary Color Additives. | Ezhini Rasendiran R, Chandresh Kumar Maurya |
| 2025 | Multilingual Query-by-Example KWS for Indian Languages using Transliteration. | Kirandevraj R, Vinod K. Kurmi, Vinay P. Namboodiri, C. V. Jawahar |
| 2025 | On the influence of language similarity in non-target speaker verification trials. | Paul M. Reuter, Michael Jessen |
| 2025 | Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model. | Yong Ren, Chenxing Li, Le Xu, Hao Gu, Duzhen Zhang, Yujie Chen, Manjie Xu, Ruibo Fu, Shan Yang, Dong Yu |
| 2025 | ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization. | Pengyu Ren, Wenhao Guan, Kaidi Wang, Peijie Chen, Qingyang Hong, Lin Li |
| 2025 | Effect of physical exercise on voice in people living with COPD. | Lauren G. Reinders, Loes van Bemmel, Alexander Mackay, David Nobbs, Frits M. E. Franssen, Hester Gietema, Simona Schfer, Sami O. Simons |
| 2025 | Focal Modulation Network: A Novel Solution for Polyphonic Music Instrument Recognition without Attention and Aggregation Strategy. | Lekshmi Chandrika Reghunath, Rajeev Rajan |
| 2025 | ASR Confidence Estimation using True Class Lexical Similarity Score. | Nagarathna Ravi, Thishyan Raj T, Ravi Teja Chaganti, Vipul Arora |
| 2025 | Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification. | William Ravenscroft, George Close, Kit Bower-Morris, Jamie Stacey, Dmitry Sityaev, Kris Y. Hong |
| 2025 | Assessing the feasibility of Large Language Models for detecting micro-behaviors in team interactions during space missions. | Ankush Raut, Projna Paromita, Sydney R. Begerowski, Suzanne T. Bell, Theodora Chaspari |
| 2025 | Synthesizing Speech with Selected Perceptual Voice Qualities - A Case Study with Creaky Voice. | Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann, Michael Kuhlmann, Petra Wagner, Reinhold Haeb-Umbach |
| 2025 | Accurate, fast, cheap: Choose three. Replacing Multi-Head-Attention with Bidirectional Recurrent Attention for Long-Form ASR. | Martin Ratajczak, Jean-Philippe Robichaud, Jennifer Drexler Fox |
| 2025 | Towards Sentence Level Imagined Speech Generation from EEG signals. | Sparsh Rastogi, Harsh Dadwal, Khushboo Modi, Jatin Bedi, Jasmeet Singh |
| 2025 | SynHate: Detecting Hate Speech in Synthetic Deepfake Audio. | Rishabh Ranjan, Kishan Pipariya, Mayank Vatsa, Richa Singh |