| 2025 | ASRU | Layer-wise Analysis for Quality of Multilingual Synthesized Speech. | Erica Cooper, Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai |
| 2025 | ASRU | PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation. | Jiajun He, Naoki Sawada, Koichi Miyazaki, Tomoki Toda |
| 2025 | ASRU | The AudioMOS Challenge 2025. | Wen-Chin Huang, Hui Wang, Cheng Liu, Yi-Chiao Wu, Andros Tjandra, Wei-Ning Hsu, Erica Cooper, Yong Qin, Tomoki Toda |
| 2025 | ASRU | Voice Factor Control Using FIR-Based Fast Neural Vocoder for Speech Generation Applications. | Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
| 2025 | ICASSP | Improvements of Discriminative Feature Space Training for Anomalous Sound Detection in Unlabeled Conditions. | Takuya Fujimura, Ibuki Kuroyanagi, Tomoki Toda |
| 2025 | ICASSP | Investigation of perceptual music similarity focusing on each instrumental part. | Yuka Hashizume, Tomoki Toda |
| 2025 | ICASSP | Investigating Factors Related to the Naturalness of Synthesized Unison Singing. | Kaito Nishizawa, Ryuichi Yamamoto, Wen-Chin Huang, Tomoki Toda |
| 2025 | ICASSP | Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual Labels. | Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai |
| 2025 | Interspeech | Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer. | Bence Mark Halpern, Thomas Tienkamp, Teja Rebernik, Rob J. J. H. van Son, Martijn Wieling, Defne Abur, Tomoki Toda |
| 2025 | Interspeech | GIA-MIC: Multimodal Emotion Recognition with Gated Interactive Attention and Modality-Invariant Learning Constraints. | Jiajun He, Jinyi Mi, Tomoki Toda |
| 2025 | Interspeech | CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models. | Jiajun He, Naoki Sawada, Koichi Miyazaki, Tomoki Toda |
| 2025 | Interspeech | SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit. | Wen-Chin Huang, Erica Cooper, Tomoki Toda |
| 2025 | Interspeech | Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition. | Cheng-Hung Hu, Yusuke Yasuda, Akifumi Yoshimoto, Tomoki Toda |
| 2025 | Interspeech | Eigenvoice Synthesis based on Model Editing for Speaker Generation. | Masato Murata, Koichi Miyazaki, Tomoki Koriyama, Tomoki Toda |
| 2025 | Interspeech | GST-BERT-TTS: Prosody Prediction Without Accentual Labels For Multi-Speaker TTS Using BERT With Global Style Tokens. | Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai |
| 2025 | Interspeech | Who, When, and What: Leveraging the "Three Ws" Concept for Emotion Recognition in Conversation. | Xiaohan Shi, Xingfeng Li, Tomoki Toda |
| 2025 | Interspeech | Speaker-Aware Multi-Task Learning for Speech Emotion Recognition. | Xiaohan Shi, Xingfeng Li, Tomoki Toda |
| 2025 | Interspeech | Advancing Emotion Recognition via Ensemble Learning: Integrating Speech, Context, and Text Representations. | Xiaohan Shi, Jinyi Mi, Xingfeng Li, Tomoki Toda |
| 2025 | Interspeech | Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments. | Reo Yoneyama, Masaya Kawamura, Ryo Terashima, Ryuichi Yamamoto, Tomoki Toda |
| 2024 | ICASSP | MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction. | Jiajun He, Xiaohan Shi, Xingfeng Li, Tomoki Toda |
| 2024 | ICASSP | Audio Difference Learning for Audio Captioning. | Tatsuya Komatsu, Yusuke Fujita, Kazuya Takeda, Tomoki Toda |
| 2024 | ICASSP | FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response Filter. | Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
| 2024 | ICASSP | Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice Conversion. | Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai |
| 2024 | ICASSP | Electrolaryngeal Speech Intelligibility Enhancement through Robust Linguistic Encoders. | Lester Phillip Violeta, Wen-Chin Huang, Ding Ma, Ryuichi Yamamoto, Kazuhiro Kobayashi, Tomoki Toda |
| 2024 | Interspeech | QHM-GAN: Neural Vocoder based on Quasi-Harmonic Modeling. | Shaowen Chen, Tomoki Toda |
| 2024 | Interspeech | Exploring the Robustness of Text-to-Speech Synthesis Based on Diffusion Probabilistic Models to Heavily Noisy Transcriptions. | Jingyi Feng, Yusuke Yasuda, Tomoki Toda |
| 2024 | Interspeech | Quantifying the effect of speech pathology on automatic and human speaker verification. | Bence Mark Halpern, Thomas Tienkamp, Wen-Chin Huang, Lester Phillip Violeta, Teja Rebernik, Sebastiaan A. H. J. de Visscher, Max J. H. Witjes, Martijn Wieling, Defne Abur, Tomoki Toda |
| 2024 | Interspeech | 2DP-2MRC: 2-Dimensional Pointer-based Machine Reading Comprehension Method for Multimodal Moment Retrieval. | Jiajun He, Tomoki Toda |
| 2024 | Interspeech | Embedding Learning for Preference-based Speech Quality Assessment. | Cheng-Hung Hu, Yusuke Yasuda, Tomoki Toda |
| 2024 | Interspeech | Challenge of Singing Voice Synthesis Using Only Text-To-Speech Corpus With FIRNet Source-Filter Neural Vocoder. | Takuma Okamoto, Yamato Ohtani, Sota Shimizu, Tomoki Toda, Hisashi Kawai |
| 2024 | Interspeech | Multimodal Fusion of Music Theory-Inspired and Self-Supervised Representations for Improved Emotion Recognition. | Xiaohan Shi, Xingfeng Li, Tomoki Toda |
| 2024 | Interspeech | CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection. | Yongyi Zang, Jiatong Shi, You Zhang, Ryuichi Yamamoto, Jionghao Han, Yuxun Tang, Shengyuan Xu, Wenxiao Zhao, Jing Guo, Tomoki Toda, Zhiyao Duan |
| 2023 | ASRU | The Voicemos Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple Domains. | Erica Cooper, Wen-Chin Huang, Yu Tsao, Hsin-Min Wang, Tomoki Toda, Junichi Yamagishi |
| 2023 | ASRU | Improving Severity Preservation of Healthy-to-Pathological Voice Conversion With Global Style Tokens. | Bence Mark Halpern, Wen-Chin Huang, Lester Phillip Violeta, R. J. J. H. van Son, Tomoki Toda |
| 2023 | ASRU | ED-CEC: Improving Rare word Recognition Using ASR Postprocessing Based on Error Detection and Context-Aware Error Correction. | Jiajun He, Zekun Yang, Tomoki Toda |
| 2023 | ASRU | The Singing Voice Conversion Challenge 2023. | Wen-Chin Huang, Lester Phillip Violeta, Songxiang Liu, Jiatong Shi, Tomoki Toda |
| 2023 | ASRU | WaveNeXt: ConvNeXt-Based Fast Neural Vocoder Without ISTFT layer. | Takuma Okamoto, Haruki Yamashita, Yamato Ohtani, Tomoki Toda, Hisashi Kawai |
| 2023 | ASRU | A Comparative Study of Voice Conversion Models With Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 2023. | Ryuichi Yamamoto, Reo Yoneyama, Lester Phillip Violeta, Wen-Chin Huang, Tomoki Toda |
| 2023 | ICASSP | Analysis Of Noisy-Target Training For Dnn-Based Speech Enhancement. | Takuya Fujimura, Tomoki Toda |
| 2023 | ICASSP | Low-Latency Electrolaryngeal Speech Enhancement Based on Fastspeech2-Based Voice Conversion and Self-Supervised Speech Representation. | Kazuhiro Kobayashi, Tomoki Hayashi, Tomoki Toda |
| 2023 | ICASSP | Representation of Vocal Tract Length Transformation Based on Group Theory. | Atsushi Miyashita, Tomoki Toda |
| 2023 | ICASSP | Intermediate Fine-Tuning Using Imperfect Synthetic Speech for Improving Electrolaryngeal Speech Recognition. | Lester Phillip Violeta, Ding Ma, Wen-Chin Huang, Tomoki Toda |
| 2023 | ICASSP | NNSVS: A Neural Network-Based Singing Voice Synthesis Toolkit. | Ryuichi Yamamoto, Reo Yoneyama, Tomoki Toda |
| 2023 | ICASSP | Text-To-Speech Synthesis Based on Latent Variable Conversion Using Diffusion Probabilistic Model and Variational Autoencoder. | Yusuke Yasuda, Tomoki Toda |
| 2023 | ICASSP | Source-Filter HiFi-GAN: Fast and Pitch Controllable High-Fidelity Neural Vocoder. | Reo Yoneyama, Yi-Chiao Wu, Tomoki Toda |
| 2023 | Interspeech | Reverberation-Controllable Voice Conversion Using Reverberation Time Estimator. | Yeonjong Choi, Chao Xie, Tomoki Toda |
| 2023 | Interspeech | Preference-based training framework for automatic speech quality assessment using deep neural network. | Cheng-Hung Hu, Yusuke Yasuda, Tomoki Toda |
| 2023 | Interspeech | E2E-S2S-VC: End-To-End Sequence-To-Sequence Voice Conversion. | Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
| 2023 | Interspeech | Emotion Awareness in Multi-utterance Turn for Improving Emotion Prediction in Multi-Speaker Conversation. | Xiaohan Shi, Xingfeng Li, Tomoki Toda |
| 2023 | Interspeech | Analysis of Mean Opinion Scores in Subjective Evaluation of Synthetic Speech Based on Tail Probabilities. | Yusuke Yasuda, Tomoki Toda |
| 2022 | ICASSP | Generalization Ability of MOS Prediction Networks. | Erica Cooper, Wen-Chin Huang, Tomoki Toda, Junichi Yamagishi |
| 2022 | ICASSP | An Investigation of Streaming Non-Autoregressive sequence-to-sequence Voice Conversion. | Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda |
| 2022 | ICASSP | LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech. | Wen-Chin Huang, Erica Cooper, Junichi Yamagishi, Tomoki Toda |
| 2022 | ICASSP | Towards Identity Preserving Normal to Dysarthric Voice Conversion. | Wen-Chin Huang, Bence Mark Halpern, Lester Phillip Violeta, Odette Scharenborg, Tomoki Toda |
| 2022 | ICASSP | S3PRL-VC: Open-Source Voice Conversion Framework with Self-Supervised Speech Representations. | Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi, Hung-Yi Lee, Shinji Watanabe, Tomoki Toda |
| 2022 | ICASSP | Direct Noisy Speech Modeling for Noisy-To-Noisy Voice Conversion. | Chao Xie, Yi-Chiao Wu, Patrick Lumban Tobing, Wen-Chin Huang, Tomoki Toda |
| 2022 | Interspeech | An Evaluation of Three-Stage Voice Conversion Framework for Noisy and Reverberant Conditions. | Yeonjong Choi, Chao Xie, Tomoki Toda |
| 2022 | Interspeech | The VoiceMOS Challenge 2022. | Wen-Chin Huang, Erica Cooper, Yu Tsao, Hsin-Min Wang, Tomoki Toda, Junichi Yamagishi |
| 2022 | Interspeech | Investigating Self-supervised Pretraining Frameworks for Pathological Speech Recognition. | Lester Phillip Violeta, Wen-Chin Huang, Tomoki Toda |
| 2022 | Interspeech | Unified Source-Filter GAN with Harmonic-plus-Noise Source Excitation Generation. | Reo Yoneyama, Yi-Chiao Wu, Tomoki Toda |
| 2022 | Interspeech | Spoken-Text-Style Transfer with Conditional Variational Autoencoder and Content Word Storage. | Daiki Yoshioka, Yusuke Yasuda, Noriyuki Matsunaga, Yamato Ohtani, Tomoki Toda |
| 2021 | ASRU | HASA-Net: A Non-Intrusive Hearing-Aid Speech Assessment Network. | Hsin-Tien Chiang, Yi-Chiao Wu, Cheng Yu, Tomoki Toda, Hsin-Min Wang, Yih-Chun Hu, Yu Tsao |
| 2021 | ASRU | On Prosody Modeling for ASR+TTS Based Voice Conversion. | Wen-Chin Huang, Tomoki Hayashi, Xinjian Li, Shinji Watanabe, Tomoki Toda |
| 2021 | ASRU | Multi-Stream HiFi-GAN with Data-Driven Waveform Decomposition. | Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
| 2021 | ASRU | Mandarin Electrolaryngeal Speech Voice Conversion with Sequence-to-Sequence Modeling. | Ming-Chi Yen, Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Shu-Wei Tsai, Yu Tsao, Tomoki Toda, Jyh-Shing Roger Jang, Hsin-Min Wang |
| 2021 | ICASSP | Speech Emotion Recognition Based on Listener Adaptive Models. | Atsushi Ando, Ryo Masumura, Hiroshi Sato, Takafumi Moriya, Takanori Ashihara, Yusuke Ijima, Tomoki Toda |
| 2021 | ICASSP | Non-Autoregressive Sequence-To-Sequence Voice Conversion. | Tomoki Hayashi, Wen-Chin Huang, Kazuhiro Kobayashi, Tomoki Toda |
| 2021 | ICASSP | Speech Recognition by Simply Fine-Tuning Bert. | Wen-Chin Huang, Chia-Hua Wu, Shang-Bao Luo, Kuan-Yu Chen, Hsin-Min Wang, Tomoki Toda |
| 2021 | ICASSP | Crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder. | Kazuhiro Kobayashi, Wen-Chin Huang, Yi-Chiao Wu, Patrick Lumban Tobing, Tomoki Hayashi, Tomoki Toda |
| 2021 | ICASSP | High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC. | Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2021 | ICASSP | Noise Level Limited Sub-Modeling for Diffusion Probabilistic Vocoders. | Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2021 | Interspeech | A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice Conversion. | Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Ching-Feng Liu, Yu Tsao, Hsin-Min Wang, Tomoki Toda |
| 2021 | Interspeech | High-Fidelity and Low-Latency Universal Neural Vocoder Based on Multiband WaveRNN with Data-Driven Linear Prediction for Discrete Waveform Modeling. | Patrick Lumban Tobing, Tomoki Toda |
| 2021 | Interspeech | Relational Data Selection for Data Augmentation of Speaker-Dependent Multi-Band MelGAN Vocoder. | Yi-Chiao Wu, Cheng-Hung Hu, Hung-Shin Lee, Yu-Huai Peng, Wen-Chin Huang, Yu Tsao, Hsin-Min Wang, Tomoki Toda |
| 2021 | Interspeech | Unified Source-Filter GAN: Unified Source-Filter Network Based On Factorization of Quasi-Periodic Parallel WaveGAN. | Reo Yoneyama, Yi-Chiao Wu, Tomoki Toda |
| 2020 | ICASSP | Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit. | Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, Xu Tan |
| 2020 | ICASSP | Weakly-Supervised Sound Event Detection with Self-Attention. | Koichi Miyazaki, Tatsuya Komatsu, Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, Kazuya Takeda |
| 2020 | ICASSP | Transformer-Based Text-to-Speech with Weighted Forced Attention. | Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2020 | ICASSP | Efficient Shallow Wavenet Vocoder Using Multiple Samples Output Based on Laplacian Distribution and Linear Prediction. | Patrick Lumban Tobing, Yi-Chiao Wu, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda |
| 2020 | Interspeech | Intelligibility Enhancement Based on Speech Waveform Modification Using Hearing Impairment. | Shu Hikosaka, Shogo Seki, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda, Hideki Banno, Tomoki Toda |
| 2020 | Interspeech | Voice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining. | Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu, Hirokazu Kameoka, Tomoki Toda |
| 2020 | Interspeech | Semi-Supervised Self-Produced Speech Enhancement and Suppression Based on Joint Source Modeling of Air- and Body-Conducted Signals Using Variational Autoencoder. | Shogo Seki, Moe Takada, Tomoki Toda |
| 2020 | Interspeech | Cyclic Spectral Modeling for Unsupervised Unit Discovery into Voice Conversion with Excitation and Waveform Modeling. | Patrick Lumban Tobing, Tomoki Hayashi, Yi-Chiao Wu, Kazuhiro Kobayashi, Tomoki Toda |
| 2020 | Interspeech | Quasi-Periodic Parallel WaveGAN Vocoder: A Non-Autoregressive Pitch-Dependent Dilated Convolution Model for Parametric Speech Generation. | Yi-Chiao Wu, Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki Toda |
| 2020 | Interspeech | A Cyclical Post-Filtering Approach to Mismatch Refinement of Neural Vocoder for Text-to-Speech Systems. | Yi-Chiao Wu, Patrick Lumban Tobing, Kazuki Yasuhara, Noriyuki Matsunaga, Yamato Ohtani, Tomoki Toda |
| 2019 | ASRU | Tacotron-Based Acoustic Model Using Phoneme Alignment for Practical Neural Text-to-Speech Systems. | Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2019 | ASRU | Investigation of Shallow Wavenet Vocoder with Laplacian Distribution Output. | Patrick Lumban Tobing, Tomoki Hayashi, Tomoki Toda |
| 2019 | ASSETS | Development of a Real-time Bionic Voice Generation System based on Statistical Excitation Prediction. | Farzaneh Ahmadi, Kazuhiro Kobayashi, Tomoki Toda |
| 2019 | ICASSP | Scene-dependent Anomalous Acoustic-event Detection Based on Conditional Wavenet and I-vector. | Tatsuya Komatsu, Tomoki Hayashi, Reishi Kondo, Tomoki Toda, Kazuya Takeda |
| 2019 | ICASSP | Investigations of Real-time Gaussian Fftnet and Parallel Wavenet Neural Vocoders with Simple Acoustic Features. | Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2019 | ICASSP | Voice Conversion with Cyclic Recurrent Neural Network and Fine-tuned Wavenet Vocoder. | Patrick Lumban Tobing, Yi-Chiao Wu, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda |
| 2019 | Interspeech | Pre-Trained Text Embeddings for Enhanced Text-to-Speech Synthesis. | Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Shubham Toshniwal, Karen Livescu |
| 2019 | Interspeech | Investigation of F0 Conditioning and Fully Convolutional Networks in Variational Autoencoder Based Voice Conversion. | Wen-Chin Huang, Yi-Chiao Wu, Chen-Chou Lo, Patrick Lumban Tobing, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda, Yu Tsao, Hsin-Min Wang |
| 2019 | Interspeech | Robustness of Statistical Voice Conversion Based on Direct Waveform Modification Against Background Sounds. | Yusuke Kurita, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
| 2019 | Interspeech | Real-Time Neural Text-to-Speech with Sequence-to-Sequence Acoustic Model and WaveGlow or Single Gaussian WaveRNN Vocoders. | Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2019 | Interspeech | Non-Parallel Voice Conversion with Cyclic Variational Autoencoder. | Patrick Lumban Tobing, Yi-Chiao Wu, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda |
| 2019 | Interspeech | Quasi-Periodic WaveNet Vocoder: A Pitch Dependent Dilated Convolution Model for Parametric Speech Generation. | Yi-Chiao Wu, Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki Toda |
| 2018 | EDUCON | Development of "KamiRepo" system with automatic student identification to handle handwritten assignments on LMS. | Shunya Seiya, Ryuya Ito, Kosuke Okamoto, Ukyo Tanikawa, Shigeki Ohira, Daisuke Deguchi, Tomoki Toda |
| 2018 | ICASSP | An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic Features. | Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2018 | ICASSP | An Investigation of Noise Shaping with Perceptual Weighting for Wavenet-Based Speech Generation. | Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2018 | Interspeech | Designing a Pneumatic Bionic Voice Prosthesis - A Statistical Approach for Source Excitation Generation. | Farzaneh Ahmadi, Tomoki Toda |
| 2018 | Interspeech | Multi-Head Decoder for End-to-End Speech Recognition. | Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, Kazuya Takeda |
| 2018 | Interspeech | Frequency Domain Variants of Velvet Noise and Their Application to Speech Processing and Synthesis. | Hideki Kawahara, Ken-Ichi Sakakibara, Masanori Morise, Hideki Banno, Tomoki Toda, Toshio Irino |
| 2018 | Interspeech | Audio-visual Voice Conversion Using Deep Canonical Correlation Analysis for Deep Bottleneck Features. | Satoshi Tamura, Kento Horio, Hajime Endo, Satoru Hayamizu, Tomoki Toda |
| 2018 | Interspeech | Collapsed Speech Segment Detection and Suppression for WaveNet Vocoder. | Yi-Chiao Wu, Kazuhiro Kobayashi, Tomoki Hayashi, Patrick Lumban Tobing, Tomoki Toda |
| 2017 | ASRU | An investigation of multi-speaker training for wavenet vocoder. | Tomoki Hayashi, Akira Tamamori, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
| 2017 | ASRU | Subband wavenet with overlapped single-sideband filterbanks. | Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2017 | ICASSP | BLSTM-HMM hybrid system combined with sound activity detection network for polyphonic Sound Event Detection. | Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, Takaaki Hori, Jonathan Le Roux, Kazuya Takeda |
| 2017 | ICASSP | A noise suppression method for body-conducted soft speech based on non-negative tensor factorization of air- and body-conducted signals. | Yusuke Tajiri, Hirokazu Kameoka, Tomoki Toda |
| 2017 | Interspeech | A Modulation Property of Time-Frequency Derivatives of Filtered Phase and its Application to Aperiodicity and f | Hideki Kawahara, Ken-Ichi Sakakibara, Masanori Morise, Hideki Banno, Tomoki Toda |
| 2017 | Interspeech | A New Cosine Series Antialiasing Function and its Application to Aliasing-Free Glottal Source Models for Speech and Singing Synthesis. | Hideki Kawahara, Ken-Ichi Sakakibara, Masanori Morise, Hideki Banno, Tomoki Toda, Toshio Irino |
| 2017 | Interspeech | Statistical Voice Conversion with WaveNet-Based Waveform Generation. | Kazuhiro Kobayashi, Tomoki Hayashi, Akira Tamamori, Tomoki Toda |
| 2017 | Interspeech | Speech Enhancement Using Non-Negative Spectrogram Models with Mel-Generalized Cepstral Regularization. | Li Li, Hirokazu Kameoka, Tomoki Toda, Shoji Makino |
| 2017 | Interspeech | Speaker-Dependent WaveNet Vocoder. | Akira Tamamori, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
| 2017 | Interspeech | Physically Constrained Statistical F | Kou Tanaka, Hirokazu Kameoka, Tomoki Toda, Satoshi Nakamura |
| 2016 | ICASSP | Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modification. | Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura |
| 2016 | ICASSP | Noise suppression method for body-conducted soft speech enhancement based on external noise monitoring. | Yusuke Tajiri, Tomoki Toda, Satoshi Nakamura |
| 2016 | ICASSP | Statistical F0 prediction for electrolaryngeal speech enhancement considering generative process of F0 contours within product of experts framework. | Kou Tanaka, Hirokazu Kameoka, Tomoki Toda, Satoshi Nakamura |
| 2016 | ICASSP | An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer. | Soichi Yamane, Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura |
| 2016 | Interspeech | A Hybrid System for Continuous Word-Level Emphasis Modeling Based on HMM State Clustering and Adaptive Training. | Quoc Truong Do, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2016 | Interspeech | The NU-NAIST Voice Conversion System for the Voice Conversion Challenge 2016. | Kazuhiro Kobayashi, Shinnosuke Takamichi, Satoshi Nakamura, Tomoki Toda |
| 2016 | Interspeech | Model Integration for HMM- and DNN-Based Speech Synthesis Using Product-of-Experts Framework. | Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2016 | Interspeech | Acoustic-to-Articulatory Inversion Mapping Based on Latent Trajectory Gaussian Mixture Model. | Patrick Lumban Tobing, Tomoki Toda, Hirokazu Kameoka, Satoshi Nakamura |
| 2016 | Interspeech | The Voice Conversion Challenge 2016. | Tomoki Toda, Ling-Hui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu, Junichi Yamagishi |
| 2015 | ACL | Improving Pivot Translation by Remembering the Pivot. | Akiva Miura, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2015 | ACL | Syntax-based Simultaneous Translation through Prediction of Unseen Syntactic Constituents. | Yusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2015 | ASRU | The NAIST ASR system for the 2015 Multi-Genre Broadcast challenge: On combination of deep learning systems using a rank-score function. | Quoc Truong Do, Michael Heck, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |
| 2015 | ASRU | A study of social-affective communication: Automatic prediction of emotion triggers and responses in television talk shows. | Nurul Lubis, Sakriani Sakti, Graham Neubig, Koichiro Yoshino, Tomoki Toda, Satoshi Nakamura |
| 2015 | ASRU | Adaptive selection from multiple response candidates in example-based dialogue. | Masahiro Mizukami, Hideaki Kizuki, Toshio Nomura, Graham Neubig, Koichiro Yoshino, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2015 | ASRU | Incremental sentence compression using LSTM recurrent networks. | Sakriani Sakti, Faiz Ilham, Graham Neubig, Tomoki Toda, Ayu Purwarianti, Satoshi Nakamura |
| 2015 | ASSETS | An Enhanced Electrolarynx with Automatic Fundamental Frequency Control based on Statistical Prediction. | Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2015 | ICASSP | EEG signal enhancement using multi-channel wiener filter with a spatial correlation prior. | Hayato Maki, Tomoki Toda, Sakriani Sakti, Graham Neubig, Satoshi Nakamura |
| 2015 | ICASSP | Parameter generation algorithm considering Modulation Spectrum for HMM-based speech synthesis. | Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura |
| 2015 | ICASSP | Modulation spectrum-constrained trajectory training algorithm for GMM-based Voice Conversion. | Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura |
| 2015 | ICASSP | Combination of two-dimensional cochleogram and spectrogram features for deep learning-based ASR. | Andros Tjandra, Sakriani Sakti, Graham Neubig, Tomoki Toda, Mirna Adriani, Satoshi Nakamura |
| 2015 | ICASSP | SAS: A speaker verification spoofing database containing diverse attacks. | Zhizheng Wu, Ali Khodabakhsh, Cenk Demiroglu, Junichi Yamagishi, Daisuke Saito, Tomoki Toda, Simon King |
| 2015 | Interspeech | Preserving word-level emphasis in speech-to-speech translation using linear regression HSMMs. | Quoc Truong Do, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |
| 2015 | Interspeech | Statistical singing voice conversion based on direct waveform modification with global variance. | Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2015 | Interspeech | Speed or accuracy? a study in evaluation of simultaneous speech translation. | Takashi Mieno, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2015 | Interspeech | A latent variable model for joint pause prediction and dependency parsing. | The Tung Nguyen, Graham Neubig, Hiroyuki Shindo, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2015 | Interspeech | Non-native speech synthesis preserving speaker individuality based on partial correction of prosodic and phonetic characteristics. | Yuji Oshima, Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2015 | Interspeech | Non-audible murmur enhancement based on statistical conversion using air- and body-conductive microphones in noisy environments. | Yusuke Tajiri, Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2015 | Interspeech | Modulation spectrum-constrained trajectory training algorithm for HMM-based speech synthesis. | Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura |
| 2015 | Interspeech | Articulatory controllable speech modification based on Gaussian mixture models with direct waveform modification using spectrum differential. | Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2015 | IUI | Automated Social Skills Trainer. | Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Tomoki Toda, Hideki Negoro, Hidemi Iwasaka, Satoshi Nakamura |
| 2015 | NAACL | Ckylark: A More Robust PCFG-LA Parser. | Yusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2015 | SP | Evaluation of a Fully Automatic Cooperative Persuasive Dialogue System. | Takuya Hiraoka, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2015 | SP | A Study on Natural Expressive Speech: Automatic Memorable Spoken Quote Detection. | Fajri Koto, Sakriani Sakti, Graham Neubig, Tomoki Toda, Mirna Adriani, Satoshi Nakamura |
| 2015 | SP | Linguistic Individuality Transformation for Spoken Language. | Masahiro Mizukami, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2015 | SP | Unknown Word Detection Based on Event-Related Brain Desynchronization Responses. | Takafumi Sasakura, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |
| 2015 | SP | An Analysis Towards Dialogue-Based Deception Detection. | Yuiko Tsunomori, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2014 | ACL | Optimizing Segmentation Strategies for Simultaneous Speech Translation. | Yusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2014 | COLING | Discriminative Language Models as a Tool for Machine Translation Error Analysis. | Koichi Akabe, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2014 | COLING | Reinforcement Learning of Cooperative Persuasive Dialogue Policies using Framing. | Takuya Hiraoka, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2014 | EACL | Acquiring a Dictionary of Emotion-Provoking Events. | Hoa Trong Vu, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2014 | ICASSP | Regression approaches to perceptual age control in singing voice conversion. | Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2014 | ICASSP | Narrow Adaptive Regularization of weights for grapheme-to-phoneme conversion. | Keigo Kubo, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |
| 2014 | ICASSP | A postfilter to modify the modulation spectrum in HMM-based speech synthesis. | Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2014 | ICASSP | An evaluation of excitation feature prediction in a hybrid approach to electrolaryngeal speech enhancement. | Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2014 | Interspeech | A hearing impairment simulation method using audiogram-based approximation of auditory charatecteristics. | Nozomi Jinbo, Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2014 | Interspeech | Excitation source analysis for high-quality speech manipulation systems based on an interference-free representation of group delay with minimum phase response compensation. | Hideki Kawahara, Masanori Morise, Tomoki Toda, Hideki Banno, Ryuichi Nisimura, Toshio Irino |
| 2014 | Interspeech | Statistical singing voice conversion with direct waveform modification based on the spectrum differential. | Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2014 | Interspeech | Structured soft margin confidence weighted learning for grapheme-to-phoneme conversion. | Keigo Kubo, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |
| 2014 | Interspeech | Data-driven generation of text balloons based on linguistic and acoustic features of a comics-anime corpus. | Sho Matsumiya, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |
| 2014 | Interspeech | Direct F | Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2014 | Interspeech | Articulatory controllable speech modification based on statistical feature mapping with Gaussian mixture models. | Patrick Lumban Tobing, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura, Ayu Purwarianti |
| 2014 | LREC | Towards Multilingual Conversations in the Medical Domain: Development of Multilingual Medical Data and A Network-based ASR System. | Sakriani Sakti, Keigo Kubo, Sho Matsumiya, Graham Neubig, Tomoki Toda, Satoshi Nakamura, Fumihiro Adachi, Ryosuke Isotani |
| 2014 | LREC | Collection of a Simultaneous Translation Corpus for Comparative Analysis. | Hiroaki Shimizu, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2013 | ACL | Towards High-Reliability Speech Translation in the Medical Domain. | Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura, Yuji Matsumoto, Ryosuke Isotani, Yukichi Ikeda |
| 2013 | ASRU | Dialogue management for leading the conversation in persuasive dialogue systems. | Takuya Hiraoka, Yuki Yamauchi, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2013 | Interspeech | Evaluation of a singing voice conversion method based on many-to-many eigenvoice conversion. | Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura |
| 2013 | Interspeech | Simple, lexicalized choice of translation timing for simultaneous speech translation. | Tomoki Fujita, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2013 | Interspeech | Generalizing continuous-space translation of paralinguistic information. | Takatomo Kano, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |
| 2013 | Interspeech | Beyond bandlimited sampling of speech spectral envelope imposed by the harmonic structure of voiced sounds. | Hideki Kawahara, Masanori Morise, Tomoki Toda, Ryuichi Nisimura, Toshio Irino |
| 2013 | Interspeech | An investigation of acoustic features for singing voice conversion based on perceptual age. | Kazuhiro Kobayashi, Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2013 | Interspeech | Grapheme-to-phoneme conversion based on adaptive regularization of weight vectors. | Keigo Kubo, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |
| 2013 | Interspeech | A digital signal processor implementation of silent/electrolaryngeal speech enhancement based on real-time statistical voice conversion. | Takuto Moriguchi, Tomoki Toda, Motoaki Sano, Hiroshi Sato, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2013 | Interspeech | An empirical comparison of joint optimization techniques for speech translation. | Masaya Ohgushi, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura |
| 2013 | Interspeech | Improvements to HMM-based speech synthesis based on parameter generation with rich context models. | Shinnosuke Takamichi, Tomoki Toda, Yoshinori Shiga, Sakriani Sakti, Graham Neubig, Satoshi Nakamura |
| 2013 | Interspeech | A hybrid approach to electrolaryngeal speech enhancement based on spectral subtraction and statistical voice conversion. | Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2012 | ICASSP | Statistical approach to voice quality control in esophageal speech enhancement. | Kenzo Yamamoto, Tomoki Toda, Hironori Doi, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2012 | Interspeech | An Evaluation of Parameter Generation Methods with Rich Context Models in HMM-Based Speech Synthesis. | Shinnosuke Takamichi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai, Sakriani Sakti, Satoshi Nakamura |
| 2012 | Interspeech | Implementation of Computationally Efficient Real-Time Voice Conversion. | Tomoki Toda, Takashi Muramatsu, Hideki Banno |
| 2011 | ASRU | Blind noise suppression for Non-Audible Murmur recognition with stereo signal processing. | Shunta Ishii, Tomoki Toda, Hiroshi Saruwatari, Sakriani Sakti, Satoshi Nakamura |
| 2011 | ICASSP | Acoustic model training for non-audible murmur recognition using transformed normal speech data. | Denis Babani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2011 | ICASSP | An evaluation of alaryngeal speech enhancement methods based on voice conversion techniques. | Hironori Doi, Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2011 | Interspeech | Speaker-Adaptive Speech Synthesis Based on Eigenvoice Conversion and Language-Dependent Prosodic Conversion in Speech-to-Speech Translation. | Nobuhiko Hattori, Tomoki Toda, Hisashi Kawai, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2010 | ICASSP | Statistical approach to enhancing esophageal speech based on Gaussian mixture models. | Hironori Doi, Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2010 | ICASSP | Non-parallel training for many-to-many eigenvoice conversion. | Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2010 | Interspeech | The use of air-pressure sensor in electrolaryngeal speech enhancement based on statistical voice conversion. | Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2010 | Interspeech | Adaptive voice-quality control based on one-to-many eigenvoice conversion. | Kumi Ohta, Tomoki Toda, Yamato Ohtani, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2010 | Interspeech | Improved training of excitation for HMM-based parametric speech synthesis. | Yoshinori Shiga, Tomoki Toda, Shinsuke Sakai, Hisashi Kawai |
| 2009 | ICASSP | Acoustic compensation methods for body transmitted speech conversion. | Daisuke Miyamoto, Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2009 | ICASSP | Voice conversion for various types of body transmitted speech. | Tomoki Toda, Keigo Nakamura, Hidehiko Sekimoto, Kiyohiro Shikano |
| 2009 | ICASSP | Trajectory training considering global variance for HMM-based speech synthesis. | Tomoki Toda, Steve J. Young |
| 2009 | ICASSP | Probablistic modelling of F0 in unvoiced regions in HMM based speech synthesis. | Kai Yu, Tomoki Toda, Milica Gasic, Simon Keizer, Franois Mairesse, Blaise Thomson, Steve J. Young |
| 2009 | Interspeech | Cross-language voice conversion based on eigenvoices. | Malorie Charlier, Yamato Ohtani, Tomoki Toda, Alexis Moinet, Thierry Dutoit |
| 2009 | Interspeech | A decision tree-based clustering approach to state definition in an excitation modeling framework for HMM-based speech synthesis. | Ranniery Maia, Tomoki Toda, Keiichi Tokuda, Shinsuke Sakai, Satoshi Nakamura |
| 2009 | Interspeech | Electrolaryngeal speech enhancement based on statistical voice conversion. | Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2009 | Interspeech | Many-to-many eigenvoice conversion with reference voice. | Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2009 | Interspeech | Technologies for processing body-conducted speech detected with non-audible murmur microphone. | Tomoki Toda, Keigo Nakamura, Takayuki Nagai, Tomomi Kaino, Yoshitaka Nakajima, Kiyohiro Shikano |
| 2009 | Interspeech | Multimodal HMM-based NAM-to-speech conversion. | Viet-Anh Tran, Grard Bailly, Hlne Loevenbruck, Tomoki Toda |
| 2008 | ICASSP | On the state definition for a trainable excitation model in HMM-based speech synthesis. | Ranniery Maia, Tomoki Toda, Keiichi Tokuda, Shinichi Sakai, Shun Nakamura |
| 2008 | ICASSP | Statistical approach to vocal tract transfer function estimation based on factor analyzed trajectory HMM. | Tomoki Toda, Keiichi Tokuda |
| 2008 | ICASSP | Performance evaluation of the speaker-independent HMM-based speech synthesis system "HTS 2007" for the Blizzard Challenge 2007. | Junichi Yamagishi, Takashi Nose, Heiga Zen, Tomoki Toda, Keiichi Tokuda |
| 2008 | Interspeech | Low-delay voice conversion based on maximum likelihood estimation of spectral parameter trajectory. | Takashi Muramatsu, Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2008 | Interspeech | Evaluation of speaking-aid system with voice conversion for laryngectomees toward its use in practical environments. | Keigo Nakamura, Tomoki Toda, Yoshitaka Nakajima, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2008 | Interspeech | An improved one-to-many eigenvoice conversion system. | Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2008 | Interspeech | Maximum a posteriori adaptation for many-to-one eigenvoice conversion. | Daisuke Tani, Tomoki Toda, Yamato Ohtani, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2008 | Interspeech | Simultaneous conversion of duration and spectrum based on statistical models including time-sequence matching. | Kaori Yutani, Yosuke Uto, Yoshihiko Nankaku, Tomoki Toda, Keiichi Tokuda |
| 2007 | ICASSP | One-to-Many and Many-to-One Voice Conversion Based on Eigenvoices. | Tomoki Toda, Yamato Ohtani, Kiyohiro Shikano |
| 2007 | Interspeech | Development of preschool children subsystem for ASR and q&a in a real-environment speech-oriented guidance task. | Tobias Cincarek, Izumi Shindo, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2007 | Interspeech | Rapid unsupervised speaker adaptation using single utterance based on MLLR and speaker selection. | Randy Gomez, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2007 | Interspeech | A trainable excitation model for HMM-based speech synthesis. | Ranniery Maia, Tomoki Toda, Heiga Zen, Yoshihiko Nankaku, Keiichi Tokuda |
| 2007 | Interspeech | Impact of various small sound source signals on voice conversion accuracy in speech communication aid for laryngectomees. | Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2007 | Interspeech | Speaker adaptive training for one-to-many eigenvoice conversion based on Gaussian mixture model. | Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2006 | ICASSP | Improving Rapid Unsupervised Speaker Adaptation Based On Hmm Sufficient Statistics. | Randy Gomez, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2006 | ICASSP | On the Use of Phonetic Information for Mapping from Articulatory Movements to Vocal Tract Spectrum. | Kenichi Nakamura, Tomoki Toda, Yoshihiko Nankaku, Keiichi Tokuda |
| 2006 | Interspeech | Acoustic modeling for spoken dialogue systems based on unsupervised utterance-based selective training. | Tobias Cincarek, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2006 | Interspeech | Improving body transmitted unvoiced speech with statistical voice conversion. | Mikihiro Nakagiri, Tomoki Toda, Hideki Kashioka, Kiyohiro Shikano |
| 2006 | Interspeech | Speaking aid system for total laryngectomees using voice conversion of body transmitted artificial speech. | Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2006 | Interspeech | Maximum likelihood voice conversion based on GMM with STRAIGHT mixed excitation. | Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2006 | Interspeech | Eigenvoice conversion based on Gaussian mixture model. | Tomoki Toda, Yamato Ohtani, Kiyohiro Shikano |
| 2006 | Interspeech | Voice conversion based on mixtures of factor analyzers. | Yosuke Uto, Yoshihiko Nankaku, Tomoki Toda, Akinobu Lee, Keiichi Tokuda |
| 2005 | ICASSP | Spectral Conversion Based on Maximum Likelihood Estimation Considering Global Variance of Converted Parameter. | Tomoki Toda, Alan W. Black, Keiichi Tokuda |
| 2005 | Interspeech | NAM-to-speech conversion with Gaussian mixture models. | Tomoki Toda, Kiyohiro Shikano |
| 2005 | Interspeech | Speech parameter generation algorithm considering global variance for HMM-based speech synthesis. | Tomoki Toda, Keiichi Tokuda |
| 2005 | Interspeech | An overview of nitech HMM-based speech synthesis system for blizzard challenge 2005. | Heiga Zen, Tomoki Toda |
| 2004 | ICASSP | An evaluation of automatic phone segmentation for concatenative speech synthesis. | Hisashi Kawai, Tomoki Toda |
| 2004 | ICASSP | Optimizing sub-cost functions for segment selection based on perceptual evaluations in concatenative speech synthesis. | Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki |
| 2004 | Interspeech | Acoustic-to-articulatory inversion mapping with Gaussian mixture model. | Tomoki Toda, Alan W. Black, Keiichi Tokuda |
| 2004 | LREC | Perceptual Evaluation of Quality Deterioration Owing to Prosody Modification. | Kazuki Adachi, Tomoki Toda, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2003 | ICASSP | Segment selection considering local degradation of naturalness in concatenative speech synthesis. | Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki, Kiyohiro Shikano |
| 2003 | Interspeech | GMM-based voice conversion applied to emotional speech synthesis. | Hiromichi Kawanami, Yohei Iwami, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2003 | Interspeech | Simple designing methods of corpus-based visual speech synthesis. | Tatsuya Shiraishi, Tomoki Toda, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2003 | Interspeech | Optimizing integrated cost function for segment selection in concatenative speech synthesis based on perceptual evaluations. | Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki |
| 2002 | ICASSP | Unit selection algorithm for Japanese speech synthesis based on both phoneme unit and diphone unit. | Tomoki Toda, Hisashi Kawai, Minoru Tsuzaki, Kiyohiro Shikano |
| 2002 | Interspeech | Designing Japanese speech database covering wide range in prosody for hybrid speech synthesizer. | Hiromichi Kawanami, Tsuyoshi Masuda, Tomoki Toda, Kiyohiro Shikano |
| 2002 | Interspeech | Evaluation of cross-language voice conversion using bilingual and non-bilingual databases. | Mikiko Mashimo, Tomoki Toda, Hiromichi Kawanami, Hideki Kashioka, Kiyohiro Shikano, Nick Campbell |
| 2002 | LREC | Designing speech database with prosodic variety for expressive TTS system. | Hiromichi Kawanami, Tsuyoshi Masuda, Tomoki Toda, Kiyohiro Shikano |
| 2001 | ICASSP | Voice conversion algorithm based on Gaussian mixture model with dynamic frequency warping of STRAIGHT spectrum. | Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2001 | Interspeech | Evaluation of cross-language voice conversion based on GMM and straight. | Mikiko Mashimo, Tomoki Toda, Kiyohiro Shikano, Nick Campbell |
| 2001 | Interspeech | High quality voice conversion based on Gaussian mixture model with dynamic frequency warping. | Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano |
| 2000 | Interspeech | Straight-based voice conversion algorithm based on Gaussian mixture model. | Tomoki Toda, Jinlin Lu, Hiroshi Saruwatari, Kiyohiro Shikano |