| 2025 | Interspeech | Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control. | Masato Murata, Koichi Miyazaki, Tomoki Koriyama |
| 2025 | Interspeech | Eigenvoice Synthesis based on Model Editing for Speaker Generation. | Masato Murata, Koichi Miyazaki, Tomoki Koriyama, Tomoki Toda |
| 2024 | Interspeech | VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features. | Tomoki Koriyama |
| 2024 | Interspeech | An Attribute Interpolation Method in Speech Synthesis by Model Merging. | Masato Murata, Koichi Miyazaki, Tomoki Koriyama |
| 2024 | Interspeech | Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech. | Dong Yang, Tomoki Koriyama, Yuki Saito |
| 2023 | ICASSP | Structured State Space Decoder for Speech Recognition and Synthesis. | Koichi Miyazaki, Masato Murata, Tomoki Koriyama |
| 2023 | ICASSP | Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-Speech. | Dong Yang, Tomoki Koriyama, Yuki Saito, Takaaki Saeki, Detai Xin, Hiroshi Saruwatari |
| 2022 | Interspeech | Predicting VQVAE-based Character Acting Style from Quotation-Annotated Text for Audiobook Speech Synthesis. | Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Yuki Saito, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
| 2022 | Interspeech | UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022. | Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Hiroshi Saruwatari |
| 2021 | Interspeech | Harmonic WaveGAN: GAN-Based Speech Waveform Generation Model with Harmonic Structure Discriminator. | Kazuki Mizuta, Tomoki Koriyama, Hiroshi Saruwatari |
| 2021 | Interspeech | Sequence-to-Sequence Learning for Deep Gaussian Process Based Speech Synthesis Using Self-Attention GP Layer. | Taiki Nakamura, Tomoki Koriyama, Hiroshi Saruwatari |
| 2021 | Interspeech | Cross-Lingual Speaker Adaptation Using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis. | Detai Xin, Yuki Saito, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
| 2020 | ICASSP | Utterance-Level Sequential Modeling for Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit. | Tomoki Koriyama, Hiroshi Saruwatari |
| 2020 | Interspeech | Multi-Speaker Text-to-Speech Synthesis Using Deep Gaussian Processes. | Kentaro Mitsui, Tomoki Koriyama, Hiroshi Saruwatari |
| 2020 | Interspeech | Cross-Lingual Text-To-Speech Synthesis via Domain Adaptation and Perceptual Similarity Regression in Speaker Space. | Detai Xin, Yuki Saito, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
| 2020 | Interspeech | Investigating Effective Additional Contextual Factors in DNN-Based Spontaneous Speech Synthesis. | Yuki Yamashita, Tomoki Koriyama, Yuki Saito, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
| 2020 | LREC | DNN-based Speech Synthesis Using Abundant Tags of Spontaneous Speech Corpus. | Yuki Yamashita, Tomoki Koriyama, Yuki Saito, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
| 2019 | ICASSP | A Training Method Using DNN-guided Layerwise Pretraining for Deep Gaussian Processes. | Tomoki Koriyama, Takao Kobayashi |
| 2019 | ICASSP | Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking. | Hiroki Tamaru, Yuki Saito, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
| 2019 | Interspeech | Semi-Supervised Prosody Modeling Using Deep Gaussian Process Latent Variable Model. | Tomoki Koriyama, Takao Kobayashi |
| 2017 | ICASSP | Duration prediction using multiple Gaussian process experts for GPR-based speech synthesis. | Decha Moungsri, Tomoki Koriyama, Takao Kobayashi |
| 2017 | Interspeech | Sampling-Based Speech Parameter Generation Using Moment-Matching Networks. | Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
| 2016 | ICASSP | A speaker adaptation technique for Gaussian process regression based speech synthesis using feature space transform. | Tomoki Koriyama, Syohei Oshio, Takao Kobayashi |
| 2016 | Interspeech | Unsupervised Stress Information Labeling Using Gaussian Process Latent Variable Model for Statistical Speech Synthesis. | Decha Moungsri, Tomoki Koriyama, Takao Kobayashi |
| 2015 | ICASSP | Prosody generation using frame-based Gaussian process regression and classification for statistical parametric speech synthesis. | Tomoki Koriyama, Takao Kobayashi |
| 2015 | Interspeech | A comparison of speech synthesis systems based on GPR, HMM, and DNN with a small amount of training data. | Tomoki Koriyama, Takao Kobayashi |
| 2015 | Interspeech | Duration prediction using multi-level model for GPR-based speech synthesis. | Decha Moungsri, Tomoki Koriyama, Takao Kobayashi |
| 2014 | ICASSP | Parametric speech synthesis based on Gaussian process regression using global variance and hyperparameter optimization. | Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
| 2014 | Interspeech | Accent type and phrase boundary estimation using acoustic and language models for automatic prosodic labeling. | Tomoki Koriyama, Hiroshi Suzuki, Takashi Nose, Takahiro Shinozaki, Takao Kobayashi |
| 2014 | Interspeech | Transform mapping using shared decision tree context clustering for HMM-based cross-lingual speech synthesis. | Daiki Nagahama, Takashi Nose, Tomoki Koriyama, Takao Kobayashi |
| 2013 | ICASSP | Frame-level acoustic modeling based on Gaussian process regression for statistical nonparametric speech synthesis. | Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
| 2013 | ICASSP | HMM-based expressive speech synthesis based on phrase-level F0 context labeling. | Yu Maeno, Takashi Nose, Takao Kobayashi, Tomoki Koriyama, Yusuke Ijima, Hideharu Nakajima, Hideyuki Mizuno, Osamu Yoshioka |
| 2013 | Interspeech | Statistical nonparametric speech synthesis using sparse Gaussian processes. | Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
| 2013 | Interspeech | A style control technique for singing voice synthesis based on multiple-regression HSMM. | Takashi Nose, Misa Kanemoto, Tomoki Koriyama, Takao Kobayashi |
| 2012 | ICASSP | An F0 modeling technique based on prosodic events for spontaneous speech synthesis. | Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
| 2012 | Interspeech | Discontinuous Observation HMM for Prosodic-Event-Based F0 Generation. | Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
| 2011 | Interspeech | On the Use of Extended Context for HMM-Based Spontaneous Conversational Speech Synthesis. | Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
| 2010 | Interspeech | Conversational spontaneous speech synthesis using average voice model. | Tomoki Koriyama, Takashi Nose, Takao Kobayashi |