| 2023 | Interspeech | eCat: An End-to-End Model for Multi-Speaker TTS & Many-to-Many Fine-Grained Prosody Transfer. | Ammar Abbas, Sri Karlapati, Bastian Schnell, Penny Karanasou, Marcel Granero Moya, Amith Nagaraj, Ayman Boustati, Nicole Peinelt, Alexis Moinet, Thomas Drugman |
| 2022 | ICASSP | Distribution Augmentation for Low-Resource Expressive Text-To-Speech. | Mateusz Lajszczak, Animesh Prasad, Arent van Korlaar, Bajibabu Bollepalli, Antonio Bonafonte, Arnaud Joly, Marco Nicolis, Alexis Moinet, Thomas Drugman, Trevor Wood, Elena Sokolova |
| 2022 | Interspeech | Expressive, Variable, and Controllable Duration Modelling in TTS. | Syed Ammar Abbas, Thomas Merritt, Alexis Moinet, Sri Karlapati, Ewa Muszynska, Simon Slangen, Elia Gatti, Thomas Drugman |
| 2022 | Interspeech | CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer. | Sri Karlapati, Penny Karanasou, Mateusz Lajszczak, Syed Ammar Abbas, Alexis Moinet, Peter Makarov, Ray Li, Arent van Korlaar, Simon Slangen, Thomas Drugman |
| 2022 | Interspeech | Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody. | Peter Makarov, Syed Ammar Abbas, Mateusz Lajszczak, Arnaud Joly, Sri Karlapati, Alexis Moinet, Thomas Drugman, Penny Karanasou |
| 2021 | ICASSP | Camp: A Two-Stage Approach to Modelling Prosody in Context. | Zack Hodari, Alexis Moinet, Sri Karlapati, Jaime Lorenzo-Trueba, Thomas Merritt, Arnaud Joly, Ammar Abbas, Penny Karanasou, Thomas Drugman |
| 2021 | ICASSP | Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech. | Sri Karlapati, Ammar Abbas, Zack Hodari, Alexis Moinet, Arnaud Joly, Penny Karanasou, Thomas Drugman |
| 2021 | ICASSP | Mispronunciation Detection in Non-Native (L2) English with Uncertainty Modeling. | Daniel Korzekwa, Jaime Lorenzo-Trueba, Szymon Zaporowski, Shira Calamaro, Thomas Drugman, Bozena Kostek |
| 2021 | Interspeech | A Learned Conditional Prior for the VAE Acoustic Space of a TTS System. | Penny Karanasou, Sri Karlapati, Alexis Moinet, Arnaud Joly, Ammar Abbas, Simon Slangen, Jaime Lorenzo-Trueba, Thomas Drugman |
| 2021 | Interspeech | Detection of Lexical Stress Errors in Non-Native (L2) English with Data Augmentation and Attention. | Daniel Korzekwa, Roberto Barra-Chicote, Szymon Zaporowski, Grzegorz Beringer, Jaime Lorenzo-Trueba, Alicja Serafinowicz, Jasha Droppo, Thomas Drugman, Bozena Kostek |
| 2021 | Interspeech | Weakly-Supervised Word-Level Pronunciation Error Detection in Non-Native English Speech. | Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman, Shira Calamaro, Bozena Kostek |
| 2020 | Interspeech | Singing Synthesis: With a Little Help from my Attention. | Orazio Angelini, Alexis Moinet, Kayoko Yanagisawa, Thomas Drugman |
| 2020 | Interspeech | CopyCat: Many-to-Many Fine-Grained Prosody Transfer for Neural Text-to-Speech. | Sri Karlapati, Alexis Moinet, Arnaud Joly, Viacheslav Klimkov, Daniel Sez-Trigueros, Thomas Drugman |
| 2020 | Interspeech | Dynamic Prosody Generation for Speech Synthesis Using Linguistics-Driven Acoustic Embedding Selection. | Shubhi Tyagi, Marco Nicolis, Jonas Rohnke, Thomas Drugman, Jaime Lorenzo-Trueba |
| 2019 | ICASSP | Effect of Data Reduction on Sequence-to-sequence Neural TTS. | Javier Latorre, Jakub Lachowicz, Jaime Lorenzo-Trueba, Thomas Merritt, Thomas Drugman, Srikanth Ronanki, Viacheslav Klimkov |
| 2019 | Interspeech | Fine-Grained Robust Prosody Transfer for Single-Speaker Neural Text-To-Speech. | Viacheslav Klimkov, Srikanth Ronanki, Jonas Rohnke, Thomas Drugman |
| 2019 | Interspeech | Interpretable Deep Learning Model for the Detection and Reconstruction of Dysarthric Speech. | Daniel Korzekwa, Roberto Barra-Chicote, Bozena Kostek, Thomas Drugman, Mateusz Lajszczak |
| 2019 | Interspeech | Towards Achieving Robust Universal Neural Vocoding. | Jaime Lorenzo-Trueba, Thomas Drugman, Javier Latorre, Thomas Merritt, Bartosz Putrycz, Roberto Barra-Chicote, Alexis Moinet, Vatsal Aggarwal |
| 2019 | NAACL | In Other News: a Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data. | Nishant Prateek, Mateusz Lajszczak, Roberto Barra-Chicote, Thomas Drugman, Jaime Lorenzo-Trueba, Thomas Merritt, Srikanth Ronanki, Trevor Wood |
| 2017 | Interspeech | Phrase Break Prediction for Long-Form Reading TTS: Exploiting Text Structure Information. | Viacheslav Klimkov, Adam Nadolski, Alexis Moinet, Bartosz Putrycz, Roberto Barra-Chicote, Thomas Merritt, Thomas Drugman |
| 2016 | Interspeech | Active and Semi-Supervised Learning in ASR: Benefits on the Acoustic and Language Models. | Thomas Drugman, Janne Pylkknen, Reinhard Kneser |
| 2016 | Interspeech | Optimizing Speech Recognition Evaluation Using Stratified Sampling. | Janne Pylkknen, Thomas Drugman, Max Bisani |
| 2015 | ICASSP | Robust excitation-based features for Automatic Speech Recognition. | Thomas Drugman, Yannis Stylianou, Langzhou Chen, Xie Chen, Mark J. F. Gales |
| 2015 | Interspeech | Fast and accurate phase unwrapping. | Thomas Drugman, Yannis Stylianou |
| 2014 | ICASSP | Parametric representation for singing voice synthesis: A comparative evaluation. | Onur Babacan, Thomas Drugman, Tuomo Raitio, Daniel Erro, Thierry Dutoit |
| 2014 | ICASSP | COVAREP - A collaborative voice analysis repository for speech technologies. | Gilles Degottex, John Kane, Thomas Drugman, Tuomo Raitio, Stefan Scherer |
| 2014 | ICASSP | Excitation modeling for HMM-based speech synthesis: Breaking down the impact of periodic and aperiodic components. | Thomas Drugman, Tuomo Raitio |
| 2014 | Interspeech | Speech synthesis in various communicative situations: impact of pronunciation variations. | Sandrine Brognaux, Benjamin Picart, Thomas Drugman |
| 2013 | ICASSP | A comparative study of pitch extraction algorithms on a large variety of singing sounds. | Onur Babacan, Thomas Drugman, Nicolas D'Alessandro, Nathalie Henrich, Thierry Dutoit |
| 2013 | ICASSP | Prediction of creaky voice from contextual factors. | Thomas Drugman, John Kane, Tuomo Raitio, Christer Gobl |
| 2013 | ICASSP | A new phase-based feature representation for robust speech recognition. | Erfan Loweimi, Seyed Mohammad Ahadi, Thomas Drugman |
| 2013 | Interspeech | A quantitative comparison of glottal closure instant estimation algorithms on a large variety of singing sounds. | Onur Babacan, Thomas Drugman, Nicolas D'Alessandro, Nathalie Henrich, Thierry Dutoit |
| 2013 | Interspeech | A new prosody annotation protocol for live sports commentaries. | Sandrine Brognaux, Benjamin Picart, Thomas Drugman |
| 2013 | Interspeech | HMM-based synthesis of creaky voice. | Tuomo Raitio, John Kane, Thomas Drugman, Christer Gobl |
| 2012 | Interspeech | Modeling the Creaky Excitation for Parametric Speech Synthesis. | Thomas Drugman, John Kane, Christer Gobl |
| 2012 | Interspeech | Resonator-based creaky voice detection. | Thomas Drugman, John Kane, Christer Gobl |
| 2012 | Interspeech | Audio and Contact Microphones for Cough Detection. | Thomas Drugman, Jrme Urbain, Nathalie Bauwens, Ricardo Chessini, Anne-Sophie Aubriot, Patrick Lebecque, Thierry Dutoit |
| 2011 | ICASSP | Phase-based information for voice pathology detection. | Thomas Drugman, Thomas Dubuisson, Thierry Dutoit |
| 2011 | Interspeech | Joint Robust Voicing Detection and Pitch Estimation Based on Residual Harmonics. | Thomas Drugman, Abeer Alwan |
| 2011 | Interspeech | Continuous Control of the Degree of Articulation in HMM-Based Speech Synthesis. | Benjamin Picart, Thomas Drugman, Thierry Dutoit |
| 2010 | Interspeech | On the potential of glottal signatures for speaker recognition. | Thomas Drugman, Thierry Dutoit |
| 2010 | Interspeech | Chirp complex cepstrum-based decomposition for asynchronous glottal analysis. | Thomas Drugman, Thierry Dutoit |
| 2010 | Interspeech | Glottal-based analysis of the lombard effect. | Thomas Drugman, Thierry Dutoit |
| 2009 | ICASSP | Using a pitch-synchronous residual codebook for hybrid HMM/frame selection speech synthesis. | Thomas Drugman, Alexis Moinet, Thierry Dutoit, Geoffrey Wilfart |
| 2009 | Interspeech | Complex cepstrum-based decomposition of speech for glottal source estimation. | Thomas Drugman, Baris Bozkurt, Thierry Dutoit |
| 2009 | Interspeech | Glottal closure and opening instant detection from speech signals. | Thomas Drugman, Thierry Dutoit |
| 2009 | Interspeech | On the mutual information between source and filter contributions for voice pathology detection. | Thomas Drugman, Thomas Dubuisson, Thierry Dutoit |
| 2009 | Interspeech | A deterministic plus stochastic model of the residual signal for improved parametric speech synthesis. | Thomas Drugman, Geoffrey Wilfart, Thierry Dutoit |
| 2008 | ICMI | Dynamic modality weighting for multi-stream hmms inaudio-visual speech recognition. | Mihai Gurban, Jean-Philippe Thiran, Thomas Drugman, Thierry Dutoit |
| 2007 | MMSP | Relevant Feature Selection for Audio-Visual Speech Recognition. | Thomas Drugman, Mihai Gurban, Jean-Philippe Thiran |