| 2025 | Interspeech | PeriodCodec: A Pitch-Controllable Neural Audio Codec Using Periodic Signals for Singing Voice Synthesis. | Masato Takagi, Miku Nishihara, Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2025 | RO-MAN | Enhancing Social Presence in Dyadic Text-Chatting with a Robot Avatar Expressing Users' Actions. | Yasutaka Nakamura, Seiichi Harata, Takuto Sakuma, Yoshihiro Tanaka, Yoshihiko Nankaku, Shohei Kato |
| 2024 | ICASSP | PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model. | Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2023 | ICASSP | Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism. | Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2023 | ICASSP | Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis System. | Takenori Yoshimura, Shinji Takaki, Kazuhiro Nakamura, Keiichiro Oura, Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2022 | ICASSP | Autoregressive Variational Autoencoder with a Hidden Semi-Markov Model-Based Structured Attention for Speech Synthesis. | Takato Fujimoto, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2022 | Interspeech | End-to-End Text-to-Speech Based on Latent Representation of Speaking Styles Using Spontaneous Dialogue. | Kentaro Mitsui, Tianyu Zhao, Kei Sawada, Yukiya Hono, Yoshihiko Nankaku, Keiichi Tokuda |
| 2021 | ICASSP | Periodnet: A Non-Autoregressive Waveform Generation Model with a Structure Separating Periodic and Aperiodic Components. | Yukiya Hono, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2020 | ICASSP | Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech Synthesis. | Takato Fujimoto, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2020 | ICASSP | Fast and High-Quality Singing Voice Synthesis System Based on Convolutional Neural Networks. | Kazuhiro Nakamura, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2020 | Interspeech | Hierarchical Multi-Grained Generative Model for Expressive Speech Synthesis. | Yukiya Hono, Kazuna Tsuboi, Kei Sawada, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2019 | ICASSP | Singing Voice Synthesis Based on Generative Adversarial Networks. | Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2019 | ICASSP | Speaker-dependent Wavenet-based Delay-free Adpcm Speech Coding. | Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2018 | ICASSP | Image Recognition Based on Separable Lattice Hmms Using a Deep Neural Network for Output Probability Distributions. | Eiji Ichikawa, Kei Sawada, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2018 | ICASSP | Statistical Voice Conversion Based on Wavenet. | Jumpei Niwa, Takenori Yoshimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2017 | ICASSP | Image recognition based on discriminative models using features generated from separable lattice HMMS. | Yoshinari Tsuzuki, Kei Sawada, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2017 | Interspeech | Articulatory Text-to-Speech Synthesis Using the Digital Waveguide Mesh Driven by a Deep Neural Network. | Amelia Jane Gully, Takenori Yoshimura, Damian T. Murphy, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2016 | ICASSP | Trajectory training considering global variance for speech synthesis based on neural networks. | Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2016 | Interspeech | Redefining the Linguistic Context Feature Set for HMM and DNN TTS Through Position and Parsing. | Rasmus Dall, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2016 | Interspeech | Voice Conversion Based on Trajectory Model Training of Neural Networks Considering Global Variance. | Naoki Hosaka, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2016 | Interspeech | Singing Voice Synthesis Based on Deep Neural Networks. | Masanari Nishimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2015 | ICASSP | The effect of neural networks in statistical parametric speech synthesis. | Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2015 | Interspeech | Prosodically-enhanced recurrent neural network language models. | Siva Reddy Gangireddy, Steve Renals, Yoshihiko Nankaku, Akinobu Lee |
| 2015 | Interspeech | Simultaneous optimization of multiple tree structures for factor analyzed HMM-based speech synthesis. | Takenori Yoshimura, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2014 | ICASSP | HMM-Based singing voice synthesis and its application to Japanese and English. | Kazuhiro Nakamura, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2014 | ICASSP | Integration of speaker and pitch adaptive training for HMM-based singing voice synthesis. | Kanako Shirota, Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2014 | Interspeech | A mel-cepstral analysis technique restoring high frequency components from low-sampling-rate speech. | Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2013 | ICASSP | Separable lattice 2-D HMMS introducing state duration control for recognition of images with various variations. | Takaya Makino, Shinji Takaki, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2013 | ICASSP | Integration of acoustic modeling and mel-cepstral analysis for HMM-based speech synthesis. | Kazuhiro Nakamura, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2013 | ICASSP | Contextual partial additive structure for HMM-based speech synthesis. | Shinji Takaki, Yoshihiko Nankaku, Keiichi Tokuda |
| 2013 | ICASSP | Image recognition based on separable lattice trajectory 2-D HMMS. | Akira Tamamori, Yoshihiko Nankaku, Keiichi Tokuda |
| 2012 | ICASSP | Face recognition based on extended separable lattice 2-D HMMS. | Keisuke Kumaki, Yoshihiko Nankaku, Keiichi Tokuda |
| 2012 | ICASSP | Pitch adaptive training for hmm-based singing voice synthesis. | Keiichiro Oura, Ayami Mase, Yoshihiko Nankaku, Keiichi Tokuda |
| 2012 | ICASSP | Face recognition based on separable lattice 2-D HMMS using variational bayesian method. | Kei Sawada, Akira Tamamori, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2012 | ICASSP | A model structure integration based on a Bayesian framework for speech recognition. | Sayaka Shiota, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2012 | Interspeech | A Bayesian Approach to Speaker Recognition Based on GMMs Using Multiple Model Structures. | Takafumi Hattori, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2012 | Interspeech | Cross-lingual Speaker Adaptation for HMM-based Speech Synthesis based on Perceptual Characteristics and Speaker Interpolation. | Viviane de Franca Oliveira, Sayaka Shiota, Yoshihiko Nankaku, Keiichi Tokuda |
| 2011 | ICASSP | Global variance modeling on frequency domain delta LSP for HMM-based speech synthesis. | Shifeng Pan, Yoshihiko Nankaku, Keiichi Tokuda, Jianhua Tao |
| 2011 | ICASSP | An optimization algorithm of independent mean and variance parameter tying structures for HMM-based speech synthesis. | Shinji Takaki, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2011 | Interspeech | Estimation of Window Coefficients for Dynamic Feature Extraction for HMM-Based Speech Synthesis. | Ling-Hui Chen, Yoshihiko Nankaku, Heiga Zen, Keiichi Tokuda, Zhen-Hua Ling, Li-Rong Dai |
| 2011 | Interspeech | Multi-Speaker Modeling with Shared Prior Distributions and Model Structures for Bayesian Speech Synthesis. | Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2011 | Interspeech | Evaluation of Tree-Trellis Based Decoding in Over-Million LVCSR. | Naoaki Ito, Yoshihiko Nankaku, Akinobu Lee |
| 2011 | Interspeech | A Bayesian Approach to Voice Conversion Based on GMMs Using Multiple Model Structures. | Lei Li, Yoshihiko Nankaku, Keiichi Tokuda |
| 2011 | Interspeech | GMM-Based Missing-Feature Reconstruction on Multi-Frame Windows. | Ulpu Remes, Yoshihiko Nankaku, Keiichi Tokuda |
| 2010 | EAMT | A Deterministic Annealing-Based Training Algorithm For Statistical Machine Translation Models. | Pascual Martnez-Gmez, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda, Germn Sanchis-Trilles |
| 2010 | ICASSP | Factor analyzed voice models for HMM-based speech synthesis. | Kyosuke Kazumi, Yoshihiko Nankaku, Keiichi Tokuda |
| 2010 | ICASSP | Face recognition based on separable lattice 2-D HMM with state duration modeling. | Yoshiaki Takahashi, Akira Tamamori, Yoshihiko Nankaku, Keiichi Tokuda |
| 2010 | ICASSP | An extension of Separable Lattice 2-D HMMS for rotational data variations. | Akira Tamamori, Yoshihiko Nankaku, Keiichi Tokuda |
| 2010 | ICASSP | Statistical parametric speech synthesis based on product of experts. | Heiga Zen, Mark J. F. Gales, Yoshihiko Nankaku, Keiichi Tokuda |
| 2010 | Interspeech | Speaker adaptation based on nonlinear spectral transform for speech recognition. | Toyohiro Hayashi, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2010 | Interspeech | HMM-based singing voice synthesis system using pitch-shifted pseudo training data. | Ayami Mase, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda |
| 2010 | Interspeech | Voice activity detection based on conditional random fields using multiple features. | Akira Saito, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2009 | ICASSP | A Bayesian approach to HMM-based speech synthesis. | Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Takashi Masuko, Keiichi Tokuda |
| 2009 | ICASSP | Voice conversion based on simultaneous modelling of spectrum and F0. | Kaori Yutani, Yosuke Uto, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2009 | ICASSP | Stereo-based stochastic noise compensation based on trajectory GMMS. | Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
| 2009 | Interspeech | A Bayesian approach to Hidden Semi-Markov Model based speech synthesis. | Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2009 | Interspeech | Tying covariance matrices to reduce the footprint of HMM-based speech synthesis systems. | Keiichiro Oura, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2009 | Interspeech | Deterministic annealing based training algorithm for Bayesian speech recognition. | Sayaka Shiota, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda |
| 2009 | Interspeech | State mapping based method for cross-lingual speaker adaptation in HMM-based speech synthesis. | Yi-Jian Wu, Yoshihiko Nankaku, Keiichi Tokuda |
| 2008 | ICASSP | Acoustic modeling with contextual additive structure for HMM-based speech recognition. | Yoshihiko Nankaku, Kazuhiro Nakamura, Heiga Zen, Keiichi Tokuda |
| 2008 | Interspeech | Bayesian context clustering using cross valid prior distribution for HMM-based speech recognition. | Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2008 | Interspeech | Speaker recognition based on variational Bayesian method. | Tatsuya Ito, Kei Hashimoto, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2008 | Interspeech | Acoustic modeling based on model structure annealing for speech recognition. | Sayaka Shiota, Kei Hashimoto, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2008 | Interspeech | Probabilistic answer selection based on conditional random fields for spoken dialog system. | Yoshitaka Yoshimi, Ryota Kakitsuba, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2008 | Interspeech | Simultaneous conversion of duration and spectrum based on statistical models including time-sequence matching. | Kaori Yutani, Yosuke Uto, Yoshihiko Nankaku, Tomoki Toda, Keiichi Tokuda |
| 2008 | Interspeech | Probabilistic feature mapping based on trajectory HMMs. | Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
| 2007 | ICASSP | Face Recognition using Hidden Markov Eigenface Models. | Yoshihiko Nankaku, Keiichi Tokuda |
| 2007 | Interspeech | A trainable excitation model for HMM-based speech synthesis. | Ranniery Maia, Tomoki Toda, Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
| 2007 | Interspeech | Model-space MLLR for trajectory HMMs. | Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
| 2006 | ICASSP | Face Recognition Based on Separable Lattice HMMS. | Daisuke Kurata, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Zoubin Ghahramani |
| 2006 | ICASSP | On the Use of Phonetic Information for Mapping from Articulatory Movements to Vocal Tract Spectrum. | Kenichi Nakamura, Tomoki Toda, Yoshihiko Nankaku, Keiichi Tokuda |
| 2006 | ICASSP | Hidden Semi-Markov Model Based Speech Recognition System using Weighted Finite-State Transducer. | Keiichiro Oura, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2006 | ICASSP | Estimating Trajectory Hmm Parameters Using Monte Carlo Em With Gibbs Sampler. | Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura |
| 2006 | Interspeech | Reducing computation on parallel decoding using frame-wise confidence scores. | Tomohiro Hakamata, Akinobu Lee, Yoshihiko Nankaku, Keiichi Tokuda |
| 2006 | Interspeech | An HMM-based singing voice synthesis system. | Keijiro Saino, Heiga Zen, Yoshihiko Nankaku, Akinobu Lee, Keiichi Tokuda |
| 2006 | Interspeech | Voice conversion based on mixtures of factor analyzers. | Yosuke Uto, Yoshihiko Nankaku, Tomoki Toda, Akinobu Lee, Keiichi Tokuda |
| 2006 | Interspeech | Speaker adaptation of trajectory HMMs using feature-space MLLR. | Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura |
| 2005 | ICASSP | Sparse KPCA for Feature Extraction in Speech Recognition. | Amaro A. de Lima, Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Fernando Gil Resende |
| 2004 | ICASSP | Parameter sharing and minimum classification error training of mixtures of factor analyzers for speaker identification. | Hiroyoshi Yamamoto, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
| 2004 | Interspeech | Deterministic annealing EM algorithm in parameter estimation for acoustic model. | Yohei Itaya, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
| 2003 | ICASSP | Speech recognition using voice-characteristic-dependent acoustic models. | Hiroyuki Suzuki, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
| 2003 | Interspeech | On the use of kernel PCA for feature extraction in speech recognition. | Amaro A. de Lima, Heiga Zen, Yoshihiko Nankaku, Chiyomi Miyajima, Keiichi Tokuda, Tadashi Kitamura |
| 2000 | ICIP | Normalized Training for HMM-Based Visual Speech Recognition. | Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Takao Kobayashi |
| 1999 | Interspeech | Intensity- and location-normalized training for HMM-based visual speech recognition. | Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura |