| 2026 | LREC | How Much Data Is Enough Data? A New Motion Capture Corpus for Probabilistic Sign Language Generation. | Anna Klezovich, Johanna Mesch, Gustav Eje Henter, Jonas Beskow |
| 2025 | ECAI | A Non-Adversarial Approach to Idempotent Generative Modelling. | Mohammed Al-Jaff, Giovanni Luca Marchetti, Michael C. Welle, Jens Lundell, Mats G. Gustafsson, Gustav Eje Henter, Hossein Azizpour, Danica Kragic |
| 2025 | HRI | Take a Look, it's in a Book, a Reading Robot. | Paige Tutts, Shivam Mehta, Zachary Syvenky, Bermet Burkanova, Mohammed Hfsafsti, Yue Wang, H. Henny Yeung, Gustav Eje Henter, Jean-Julien Aucouturier, Angelica Lim |
| 2025 | Interspeech | SawtArabi: A Benchmark Corpus for Arabic TTS. Standard, Dialectal and Code-Switching. | Vasista Sai Lodagala, Lamya Alkanhal, Daniel Izham, Shivam Mehta, Shammur Absar Chowdhury, Aqeelah Makki, Hamdy S. Hussein, Gustav Eje Henter, Ahmed Ali |
| 2025 | Interspeech | Hear Me Out: Interactive evaluation and bias discovery platform for speech-to-speech conversational AI. | Shree Harsha Bokkahalli Satish, Gustav Eje Henter, va Szkely |
| 2025 | RO-MAN | EmojiVoice: Towards long-term controllable expressivity in robot speech. | Paige Tutts, Shivam Mehta, Zachary Syvenky, Bermet Burkanova, Gustav Eje Henter, Angelica Lim |
| 2024 | CVPR | Fake it to make it: Using synthetic data to remedy the data shortage in joint multi-modal speech-and-gesture synthesis. | Shivam Mehta, Anna Deichler, Jim O'Regan, Birger Moll, Jonas Beskow, Gustav Eje Henter, Simon Alexanderson |
| 2024 | ICASSP | Unified Speech and Gesture Synthesis Using Flow Matching. | Shivam Mehta, Ruibo Tu, Simon Alexanderson, Jonas Beskow, va Szkely, Gustav Eje Henter |
| 2024 | ICASSP | Matcha-TTS: A Fast TTS Architecture with Conditional Flow Matching. | Shivam Mehta, Ruibo Tu, Jonas Beskow, va Szkely, Gustav Eje Henter |
| 2024 | ICMI | GENEA Workshop 2024: The 5th Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied Agents. | Youngwoo Yoon, Taras Kucherenko, Alice Delbosc, Rajmund Nagy, Teodor Nikolov, Gustav Eje Henter |
| 2024 | Interspeech | Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech. | Shivam Mehta, Harm Lameris, Rajiv Punmiya, Jonas Beskow, va Szkely, Gustav Eje Henter |
| 2023 | ICASSP | Prosody-Controllable Spontaneous TTS with Neural HMMS. | Harm Lameris, Shivam Mehta, Gustav Eje Henter, Joakim Gustafson, va Szkely |
| 2023 | ICASSP | A Comparative Study of Self-Supervised Speech Representations in Read and Spontaneous TTS. | Siyang Wang, Gustav Eje Henter, Joakim Gustafson, va Szkely |
| 2023 | ICASSP | Autovocoder: Fast Waveform Generation from a Learned Speech Representation Using Differentiable Digital Signal Processing. | Jacob J. Webber, Cassia Valentini-Botinhao, Evelyn Williams, Gustav Eje Henter, Simon King |
| 2023 | ICASSP | A Processing Framework to Access Large Quantities of Whispered Speech Found in ASMR. | Pablo Prez Zarazaga, Gustav Eje Henter, Zofia Malisz |
| 2023 | ICMI | The GENEA Challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings. | Taras Kucherenko, Rajmund Nagy, Youngwoo Yoon, Jieyeon Woo, Teodor Nikolov, Mihail Tsakov, Gustav Eje Henter |
| 2023 | ICMI | "Am I listening?", Evaluating the Quality of Generated Data-driven Listening Motion. | Pieter Wolfert, Gustav Eje Henter, Tony Belpaeme |
| 2023 | ICMI | GENEA Workshop 2023: The 4th Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied Agents. | Youngwoo Yoon, Taras Kucherenko, Jieyeon Woo, Pieter Wolfert, Rajmund Nagy, Gustav Eje Henter |
| 2023 | Interspeech | OverFlow: Putting flows on top of neural transducers for better TTS. | Shivam Mehta, Ambika Kirkland, Harm Lameris, Jonas Beskow, va Szkely, Gustav Eje Henter |
| 2023 | Interspeech | Speaker-independent neural formant synthesis. | Pablo Prez Zarazaga, Zofia Malisz, Gustav Eje Henter, Lauri Juvela |
| 2022 | ICASSP | Wavebender GAN: An Architecture for Phonetically Meaningful Speech Manipulation. | Gustavo Teodoro Dhler Beck, Ulme Wennberg, Zofia Malisz, Gustav Eje Henter |
| 2022 | ICASSP | Neural HMMS Are All You Need (For High-Quality Attention-Free TTS). | Shivam Mehta, va Szkely, Jonas Beskow, Gustav Eje Henter |
| 2022 | ICMI | GENEA Workshop 2022: The 3rd Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied Agents. | Pieter Wolfert, Taras Kucherenko, Carla Viegas, Zerrin Yumak, Youngwoo Yoon, Gustav Eje Henter |
| 2022 | ICMI | The GENEA Challenge 2022: A large evaluation of data-driven co-speech gesture generation. | Youngwoo Yoon, Pieter Wolfert, Taras Kucherenko, Carla Viegas, Teodor Nikolov, Mihail Tsakov, Gustav Eje Henter |
| 2022 | Interspeech | Speech Audio Corrector: using speech from non-target speakers for one-off correction of mispronunciations in grapheme-input text-to-speech. | Jason Fong, Daniel Lyth, Gustav Eje Henter, Hao Tang, Simon King |
| 2022 | Interspeech | Predicting pairwise preferences between TTS audio stimuli using parallel ratings data and anti-symmetric twin neural networks. | Cassia Valentini-Botinhao, Manuel Sam Ribeiro, Oliver Watts, Korin Richmond, Gustav Eje Henter |
| 2021 | ACL | The Case for Translation-Invariant Self-Attention in Transformer-Based Language Models. | Ulme Wennberg, Gustav Eje Henter |
| 2021 | ICMI | HEMVIP: Human Evaluation of Multiple Videos in Parallel. | Patrik Jonell, Youngwoo Yoon, Pieter Wolfert, Taras Kucherenko, Gustav Eje Henter |
| 2021 | ICMI | GENEA Workshop 2021: The 2nd Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied Agents. | Taras Kucherenko, Patrik Jonell, Youngwoo Yoon, Pieter Wolfert, Zerrin Yumak, Gustav Eje Henter |
| 2021 | ICMI | Integrated Speech and Gesture Synthesis. | Siyang Wang, Simon Alexanderson, Joakim Gustafson, Jonas Beskow, Gustav Eje Henter, va Szkely |
| 2021 | IUI | A Large, Crowdsourced Evaluation of Gesture Generation Systems on Common Data: The GENEA Challenge 2020. | Taras Kucherenko, Patrik Jonell, Youngwoo Yoon, Pieter Wolfert, Gustav Eje Henter |
| 2021 | IVA | Speech2Properties2Gestures: Gesture-Property Prediction as a Tool for Generating Representational Gestures from Speech. | Taras Kucherenko, Rajmund Nagy, Patrik Jonell, Michael Neff, Hedvig Kjellstrm, Gustav Eje Henter |
| 2020 | ICASSP | Breathing and Speech Planning in Spontaneous Speech Synthesis. | va Szkely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson |
| 2020 | ICMI | Gesticulator: A framework for semantically-aware speech-driven gesture generation. | Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexandersson, Iolanda Leite, Hedvig Kjellstrm |
| 2020 | IVA | Generating coherent spontaneous speech and gesture from text. | Simon Alexanderson, va Szkely, Gustav Eje Henter, Taras Kucherenko, Jonas Beskow |
| 2020 | IVA | Let's Face It: Probabilistic Multi-modal Interlocutor-aware Generation of Facial Gestures in Dyadic Settings. | Patrik Jonell, Taras Kucherenko, Gustav Eje Henter, Jonas Beskow |
| 2019 | ICASSP | Casting to Corpus: Segmenting and Selecting Spontaneous Dialogue for Tts with a Cnn-lstm Speaker-dependent Breath Detector. | va Szkely, Gustav Eje Henter, Joakim Gustafson |
| 2019 | Interspeech | Off the Cuff: Exploring Extemporaneous Speech Delivery with TTS. | va Szkely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson |
| 2019 | Interspeech | Spontaneous Conversational Speech Synthesis from Found Data. | va Szkely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson |
| 2019 | IVA | Analyzing Input and Output Representations for Speech-Driven Gesture Generation. | Taras Kucherenko, Dai Hasegawa, Gustav Eje Henter, Naoshi Kaneko, Hedvig Kjellstrm |
| 2018 | ICASSP | Cyborg Speech: Deep Multilingual Speech Synthesis for Generating Segmental Foreign Accent with Natural Prosody. | Gustav Eje Henter, Jaime Lorenzo-Trueba, Xin Wang, Mariko Kondo, Junichi Yamagishi |
| 2017 | ICASSP | Adapting and controlling DNN-based speech synthesis using input codes. | Hieu-Thi Luong, Shinji Takaki, Gustav Eje Henter, Junichi Yamagishi |
| 2017 | Interspeech | Principles for Learning Controllable TTS from Annotated and Latent Variation. | Gustav Eje Henter, Jaime Lorenzo-Trueba, Xin Wang, Junichi Yamagishi |
| 2017 | Interspeech | Misperceptions of the Emotional Content of Natural and Vocoded Speech in a Car. | Jaime Lorenzo-Trueba, Cassia Valentini-Botinhao, Gustav Eje Henter, Junichi Yamagishi |
| 2016 | ICASSP | Testing the consistency assumption: Pronunciation variant forced alignment in read and spontaneous speech synthesis. | Rasmus Dall, Sandrine Brognaux, Korin Richmond, Cassia Valentini-Botinhao, Gustav Eje Henter, Julia Hirschberg, Junichi Yamagishi, Simon King |
| 2016 | ICASSP | Robust TTS duration modelling using DNNS. | Gustav Eje Henter, Srikanth Ronanki, Oliver Watts, Mirjam Wester, Zhizheng Wu, Simon King |
| 2016 | ICASSP | From HMMS to DNNS: Where do the improvements come from? | Oliver Watts, Gustav Eje Henter, Thomas Merritt, Zhizheng Wu, Simon King |
| 2016 | Interspeech | A Template-Based Approach for Speech Synthesis Intonation Generation Using LSTMs. | Srikanth Ronanki, Gustav Eje Henter, Zhizheng Wu, Simon King |
| 2016 | Interspeech | A Hierarchical Predictor of Synthetic Speech Naturalness Using Neural Networks. | Takenori Yoshimura, Gustav Eje Henter, Oliver Watts, Mirjam Wester, Junichi Yamagishi, Keiichi Tokuda |
| 2015 | Interspeech | Are we using enough listeners? no! - an empirically-supported critique of interspeech 2014 TTS evaluations. | Mirjam Wester, Cassia Valentini-Botinhao, Gustav Eje Henter |
| 2014 | Interspeech | A flexible front-end for HTS. | Matthew P. Aylett, Rasmus Dall, Arnab Ghoshal, Gustav Eje Henter, Thomas Merritt |
| 2014 | Interspeech | Measuring the perceptual effects of modelling assumptions in speech synthesis using stimuli constructed from repeated natural speech. | Gustav Eje Henter, Thomas Merritt, Matt Shannon, Catherine Mayo, Simon King |
| 2012 | ICASSP | Gaussian process dynamical models for nonparametric speech representation and synthesis. | Gustav Eje Henter, Marcus R. Frean, W. Bastiaan Kleijn |
| 2012 | Interspeech | Enhancing Subjective Speech Intelligibility Using a Statistical Model of Speech. | Petko Nikolov Petkov, W. Bastiaan Kleijn, Gustav Eje Henter |
| 2011 | Interspeech | Intermediate-State HMMs to Capture Continuously-Changing Signal Features. | Gustav Eje Henter, W. Bastiaan Kleijn |