| 2026 | CHI | From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction. | Shree Harsha Bokkahalli Satish, Maria Teleki, Christoph Minixhofer, Ondrej Klejch, Peter Bell, va Szkely |
| 2026 | IUI | "Walk a Mile in My Voice": Voice Conversion Shapes Trust, Attribution, and Empathy in Human-AI Speech Interactions. | Shree Harsha Bokkahalli Satish, Maria Teleki, Christoph Minixhofer, Ondrej Klejch, Peter Bell, va Szkely |
| 2025 | Interspeech | From Static to Dynamic: Enhancing AAC with Generative Imagery and Zero-Shot TTS. | Juliana Francis, Joakim Gustafson, va Szkely |
| 2025 | Interspeech | Voices of 'cyborg awesomeness': Posthuman embodiment of nonbinary gender expression in AI speech technologies. | Maxwell Hope, va Szkely |
| 2025 | Interspeech | VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech. | Harm Lameris, Joakim Gustafson, va Szkely |
| 2025 | Interspeech | Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM. | Dariia Puhach, Amir H. Payberah, va Szkely |
| 2025 | Interspeech | Hear Me Out: Interactive evaluation and bias discovery platform for speech-to-speech conversational AI. | Shree Harsha Bokkahalli Satish, Gustav Eje Henter, va Szkely |
| 2025 | Interspeech | Voice Reconstruction through Large-Scale TTS Models: Comparing Zero-Shot and Fine-tuning Approaches to Personalise TTS in Assistive Communication. | va Szkely, Pter Mihajlik, Mt Soma Kdr, Lszl Tth |
| 2024 | COLING | The Role of Creaky Voice in Turn Taking and the Perception of Speaker Stance: Experiments Using Controllable TTS. | Harm Lameris, va Szkely, Joakim Gustafson |
| 2024 | COLING | Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model. | Siyang Wang, va Szkely |
| 2024 | ICASSP | Unified Speech and Gesture Synthesis Using Flow Matching. | Shivam Mehta, Ruibo Tu, Simon Alexanderson, Jonas Beskow, va Szkely, Gustav Eje Henter |
| 2024 | ICASSP | Matcha-TTS: A Fast TTS Architecture with Conditional Flow Matching. | Shivam Mehta, Ruibo Tu, Jonas Beskow, va Szkely, Gustav Eje Henter |
| 2024 | Interspeech | ConnecTone: a modular AAC system prototype with contextual generative text prediction and style-adaptive conversational TTS. | Juliana Francis, va Szkely, Joakim Gustafson |
| 2024 | Interspeech | CreakVC: a voice conversion tool for modulating creaky voice. | Harm Lameris, Joakim Gustafson, va Szkely |
| 2024 | Interspeech | Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech. | Shivam Mehta, Harm Lameris, Rajiv Punmiya, Jonas Beskow, va Szkely, Gustav Eje Henter |
| 2024 | Interspeech | Well, what can you do with messy data? Exploring the prosody and pragmatic function of the discourse marker "well" with found data and speech synthesis. | Johannah O'Mahony, Catherine Lai, va Szkely |
| 2024 | Interspeech | An inclusive approach to creating a palette of synthetic voices for gender diversity. | va Szkely, Maxwell Hope |
| 2024 | Interspeech | Contextual Interactive Evaluation of TTS Models in Dialogue Systems. | Siyang Wang, va Szkely, Joakim Gustafson |
| 2024 | SIGdial | Voice and Choice: Investigating the Role of Prosodic Variation in Request Compliance and Perceived Politeness Using Conversational TTS. | va Szkely, Jeff Higginbotham, Francesco Possemato |
| 2023 | HAI | Why is my Agent so Slow? Deploying Human-Like Conversational Turn-Taking. | Matthew Peter Aylett, va Szkely, Donald McMillan, Gabriel Skantze, Marta Romeo, Joel E. Fischer, Gisela Reyes-Cruz |
| 2023 | ICASSP | Prosody-Controllable Spontaneous TTS with Neural HMMS. | Harm Lameris, Shivam Mehta, Gustav Eje Henter, Joakim Gustafson, va Szkely |
| 2023 | ICASSP | A Comparative Study of Self-Supervised Speech Representations in Read and Spontaneous TTS. | Siyang Wang, Gustav Eje Henter, Joakim Gustafson, va Szkely |
| 2023 | Interspeech | Automatic Evaluation of Turn-taking Cues in Conversational Speech Synthesis. | Erik Ekstedt, Siyang Wang, va Szkely, Joakim Gustafson, Gabriel Skantze |
| 2023 | Interspeech | Synthesis after a couple PINTs: Investigating the Role of Pause-Internal Phonetic Particles in Speech Synthesis and Perception. | Mikey Elmers, Johannah O'Mahony, va Szkely |
| 2023 | Interspeech | Pardon my disfluency: The impact of disfluency effects on the perception of speaker competence and confidence. | Ambika Kirkland, Joakim Gustafson, va Szkely |
| 2023 | Interspeech | Beyond Style: Synthesizing Speech with Pragmatic Functions. | Harm Lameris, Joakim Gustafson, va Szkely |
| 2023 | Interspeech | OverFlow: Putting flows on top of neural transducers for better TTS. | Shivam Mehta, Ambika Kirkland, Harm Lameris, Jonas Beskow, va Szkely, Gustav Eje Henter |
| 2023 | Interspeech | Prosody-controllable Gender-ambiguous Speech Synthesis: A Tool for Investigating Implicit Bias in Speech Perception. | va Szkely, Joakim Gustafson, Ilaria Torre |
| 2023 | Interspeech | So-to-Speak: An Exploratory Platform for Investigating the Interplay between Style and Prosody in TTS. | va Szkely, Siyang Wang, Joakim Gustafson |
| 2023 | IVA | Generation of speech and facial animation with controllable articulatory effort for amusing conversational characters. | Joakim Gustafson, va Szkely, Jonas Beskow |
| 2023 | RO-MAN | Hi robot, it's not what you say, it's how you say it. | Jura Miniota, Siyang Wang, Jonas Beskow, Joakim Gustafson, va Szkely, Andr Pereira |
| 2023 | RO-MAN | Can a gender-ambiguous voice reduce gender stereotypes in human-robot interactions? | Ilaria Torre, Erik Lagerstedt, Nathaniel Dennler, Katie Seaborn, Iolanda Leite, va Szkely |
| 2022 | ICASSP | Neural HMMS Are All You Need (For High-Quality Attention-Free TTS). | Shivam Mehta, va Szkely, Jonas Beskow, Gustav Eje Henter |
| 2022 | Interspeech | Where's the uh, hesitation? The interplay between filled pause location, speech rate and fundamental frequency in perception of confidence. | Ambika Kirkland, Harm Lameris, va Szkely, Joakim Gustafson |
| 2022 | LREC | Evaluating Sampling-based Filler Insertion with Spontaneous TTS. | Siyang Wang, Joakim Gustafson, va Szkely |
| 2021 | ICMI | Integrated Speech and Gesture Synthesis. | Siyang Wang, Simon Alexanderson, Joakim Gustafson, Jonas Beskow, Gustav Eje Henter, va Szkely |
| 2020 | ICASSP | Breathing and Speech Planning in Spontaneous Speech Synthesis. | va Szkely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson |
| 2020 | IVA | Generating coherent spontaneous speech and gesture from text. | Simon Alexanderson, va Szkely, Gustav Eje Henter, Taras Kucherenko, Jonas Beskow |
| 2020 | LREC | Augmented Prompt Selection for Evaluation of Spontaneous Speech Synthesis. | va Szkely, Jens Edlund, Joakim Gustafson |
| 2019 | CHI | Mapping Theoretical and Methodological Perspectives for Understanding Speech Interface Interactions. | Leigh Clark, Benjamin R. Cowan, Justin Edwards, Cosmin Munteanu, Christine Murad, Matthew P. Aylett, Roger K. Moore, Jens Edlund, va Szkely, Patrick Healey, Naomi Harte, Ilaria Torre, Philip R. Doyle |
| 2019 | ICASSP | Casting to Corpus: Segmenting and Selecting Spontaneous Dialogue for Tts with a Cnn-lstm Speaker-dependent Breath Detector. | va Szkely, Gustav Eje Henter, Joakim Gustafson |
| 2019 | Interspeech | The Greennn Tree - Lengthening Position Influences Uncertainty Perception. | Simon Betz, Sina Zarrie, va Szkely, Petra Wagner |
| 2019 | Interspeech | Off the Cuff: Exploring Extemporaneous Speech Delivery with TTS. | va Szkely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson |
| 2019 | Interspeech | Spontaneous Conversational Speech Synthesis from Found Data. | va Szkely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson |
| 2017 | CogSci | They Know as Much as We Do: Knowledge Estimation and Partner Modelling of Artificial Partners. | Benjamin R. Cowan, Holly P. Branigan, Habiba Begum, Lucy McKenna, va Szkely |
| 2017 | ICMI | Using crowd-sourcing for the design of listening agents: challenges and opportunities. | Catharine Oertel, Patrik Jonell, Kevin El Haddad, va Szkely, Joakim Gustafson |
| 2017 | Interspeech | Synthesising Uncertainty: The Interplay of Vocal Effort and Hesitation Disfluencies. | va Szkely, Joseph Mendelson, Joakim Gustafson |
| 2015 | Interspeech | The effect of soft, modal and loud voice levels on entrainment in noisy conditions. | va Szkely, Mark T. Keane, Julie Carson-Berndsen |
| 2013 | IUI | A system for facial expression-based affective speech translation. | Zeeshan Ahmed, Ingmar Steiner, va Szkely, Julie Carson-Berndsen |
| 2012 | ICASSP | Detecting a targeted voice style in an audiobook using voice quality features. | va Szkely, John Kane, Stefan Scherer, Christer Gobl, Julie Carson-Berndsen |
| 2012 | LREC | Rapidly Testing the Interaction Model of a Pronunciation Training System via Wizard-of-Oz. | Joo P. Cabral, Mark Kane, Zeeshan Ahmed, Mohamed Abou-Zleikha, va Szkely, Amalia Zahra, Kalu U. Ogbureke, Peter Cahill, Julie Carson-Berndsen, Stephan Schlgl |
| 2012 | LREC | Evaluating expressive speech synthesis from audiobook corpora for conversational phrases. | va Szkely, Joo P. Cabral, Mohamed Abou-Zleikha, Peter Cahill, Julie Carson-Berndsen |
| 2011 | Interspeech | Clustering Expressive Speech Styles in Audiobooks Using Glottal Source Parameters. | va Szkely, Joo P. Cabral, Peter Cahill, Julie Carson-Berndsen |