| 2025 | ICASSP | Lightweight neural front-ends for low-resource on-device Text-to-Speech. | Giulia Comini, Heereen Shim, Manuel Sam Ribeiro |
| 2023 | Interspeech | Multilingual context-based pronunciation learning for Text-to-Speech. | Giulia Comini, Manuel Sam Ribeiro, Fan Yang, Heereen Shim, Jaime Lorenzo-Trueba |
| 2023 | Interspeech | Improving grapheme-to-phoneme conversion by learning pronunciations from speech recordings. | Manuel Sam Ribeiro, Giulia Comini, Jaime Lorenzo-Trueba |
| 2023 | Interspeech | Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech. | Guangyan Zhang, Thomas Merritt, Manuel Sam Ribeiro, Biel Tura Vecino, Kayoko Yanagisawa, Kamil Pokora, Abdelhamid Ezzerg, Sebastian Cygert, Ammar Abbas, Piotr Bilinski, Roberto Barra-Chicote, Daniel Korzekwa, Jaime Lorenzo-Trueba |
| 2022 | ICASSP | Voice Filter: Few-Shot Text-to-Speech Speaker Adaptation Using Voice Conversion as a Post-Processing Module. | Adam Gabrys, Goeric Huybrechts, Manuel Sam Ribeiro, Chung-Ming Chien, Julian Roth, Giulia Comini, Roberto Barra-Chicote, Bartek Perz, Jaime Lorenzo-Trueba |
| 2022 | ICASSP | Cross-Speaker Style Transfer for Text-to-Speech Using Data Augmentation. | Manuel Sam Ribeiro, Julian Roth, Giulia Comini, Goeric Huybrechts, Adam Gabrys, Jaime Lorenzo-Trueba |
| 2022 | Interspeech | Low-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentation. | Giulia Comini, Goeric Huybrechts, Manuel Sam Ribeiro, Adam Gabrys, Jaime Lorenzo-Trueba |
| 2022 | Interspeech | Predicting pairwise preferences between TTS audio stimuli using parallel ratings data and anti-symmetric twin neural networks. | Cassia Valentini-Botinhao, Manuel Sam Ribeiro, Oliver Watts, Korin Richmond, Gustav Eje Henter |
| 2021 | Interspeech | Silent versus Modal Multi-Speaker Speech Recognition from Ultrasound and Video. | Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals |
| 2019 | ICASSP | Speaker-independent Classification of Phonetic Segments from Raw Ultrasound in Child Speech. | Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals |
| 2019 | Interspeech | Synchronising Audio and Ultrasound by Learning Cross-Modal Embeddings. | Aciel Eshky, Manuel Sam Ribeiro, Korin Richmond, Steve Renals |
| 2019 | Interspeech | Ultrasound Tongue Imaging for Diarization and Alignment of Child Speech Therapy Sessions. | Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals |
| 2018 | Interspeech | UltraSuite: A Repository of Ultrasound and Acoustic Data from Child Speech Therapy Sessions. | Aciel Eshky, Manuel Sam Ribeiro, Joanne Cleland, Korin Richmond, Zoe Roxburgh, James M. Scobbie, Alan Wrench |
| 2017 | Interspeech | Learning Word Vector Representations Based on Acoustic Counts. | Manuel Sam Ribeiro, Oliver Watts, Junichi Yamagishi |
| 2016 | ICASSP | Wavelet-based decomposition of F0 as a secondary task for DNN-based speech synthesis with multi-task learning. | Manuel Sam Ribeiro, Oliver Watts, Junichi Yamagishi, Robert A. J. Clark |
| 2016 | Interspeech | The SIWIS Database: A Multilingual Speech Database with Acted Emphasis. | Jean-Philippe Goldman, Pierre-Edouard Honnet, Robert A. J. Clark, Philip N. Garner, Maria Ivanova, Alexandros Lazaridis, Hui Liang, Tiago Macedo, Beat Pfister, Manuel Sam Ribeiro, Eric Wehrli, Junichi Yamagishi |
| 2016 | Interspeech | Syllable-Level Representations of Suprasegmental Features for DNN-Based Text-to-Speech Synthesis. | Manuel Sam Ribeiro, Oliver Watts, Junichi Yamagishi |
| 2015 | ICASSP | A multi-level representation of f0 using the continuous wavelet transform and the Discrete Cosine Transform. | Manuel Sam Ribeiro, Robert A. J. Clark |
| 2015 | Interspeech | A perceptual investigation of wavelet-based decomposition of f0 for text-to-speech synthesis. | Manuel Sam Ribeiro, Junichi Yamagishi, Robert A. J. Clark |