| 2025 | ICASSP | Investigating Factors Related to the Naturalness of Synthesized Unison Singing. | Kaito Nishizawa, Ryuichi Yamamoto, Wen-Chin Huang, Tomoki Toda |
| 2025 | ICASSP | Description-Based Controllable Text-to-Speech With Cross-Lingual Voice Control. | Ryuichi Yamamoto, Yuma Shirahata, Masaya Kawamura, Kentaro Tachibana |
| 2025 | Interspeech | BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing. | Masaya Kawamura, Takuya Hasumi, Yuma Shirahata, Ryuichi Yamamoto |
| 2025 | Interspeech | Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning. | Hien Ohnaka, Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto |
| 2025 | Interspeech | Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments. | Reo Yoneyama, Masaya Kawamura, Ryo Terashima, Ryuichi Yamamoto, Tomoki Toda |
| 2024 | ICASSP | PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language Descriptions. | Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, Kentaro Tachibana |
| 2024 | ICASSP | Electrolaryngeal Speech Intelligibility Enhancement through Robust Linguistic Encoders. | Lester Phillip Violeta, Wen-Chin Huang, Ding Ma, Ryuichi Yamamoto, Kazuhiro Kobayashi, Tomoki Toda |
| 2024 | ICASSP | Enhancing Multilingual TTS with Voice Conversion Based Data Augmentation and Posterior Embedding. | Hyun-Wook Yoon, Jin-Seob Kim, Ryuichi Yamamoto, Ryo Terashima, Chan-Ho Song, Jae-Min Kim, Eunwoo Song |
| 2024 | Interspeech | Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment. | Takuto Igarashi, Yuki Saito, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari |
| 2024 | Interspeech | LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning. | Masaya Kawamura, Ryuichi Yamamoto, Yuma Shirahata, Takuya Hasumi, Kentaro Tachibana |
| 2024 | Interspeech | SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark. | Yuki Saito, Takuto Igarashi, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari |
| 2024 | Interspeech | Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data. | Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
| 2024 | Interspeech | CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection. | Yongyi Zang, Jiatong Shi, You Zhang, Ryuichi Yamamoto, Jionghao Han, Yuxun Tang, Shengyuan Xu, Wenxiao Zhao, Jing Guo, Tomoki Toda, Zhiyao Duan |
| 2023 | ASRU | A Comparative Study of Voice Conversion Models With Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 2023. | Ryuichi Yamamoto, Reo Yoneyama, Lester Phillip Violeta, Wen-Chin Huang, Tomoki Toda |
| 2023 | ICASSP | Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform. | Masaya Kawamura, Yuma Shirahata, Ryuichi Yamamoto, Kentaro Tachibana |
| 2023 | ICASSP | Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech Synthesis. | Yuma Shirahata, Ryuichi Yamamoto, Eunwoo Song, Ryo Terashima, Jae-Min Kim, Kentaro Tachibana |
| 2023 | ICASSP | NNSVS: A Neural Network-Based Singing Voice Synthesis Toolkit. | Ryuichi Yamamoto, Reo Yoneyama, Tomoki Toda |
| 2023 | ICASSP | Nonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs. | Reo Yoneyama, Ryuichi Yamamoto, Kentaro Tachibana |
| 2022 | Interspeech | A Unified Accent Estimation Method Based on Multi-Task Learning for Japanese Text-to-Speech. | Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
| 2022 | Interspeech | DRSpeech: Degradation-Robust Text-to-Speech Synthesis with Frame-Level and Utterance-Level Acoustic Representation Learning. | Takaaki Saeki, Kentaro Tachibana, Ryuichi Yamamoto |
| 2022 | Interspeech | TTS-by-TTS 2: Data-Selective Augmentation for Neural Speech Synthesis Using Ranking Support Vector Machine with Variational Autoencoder. | Eunwoo Song, Ryuichi Yamamoto, Ohsung Kwon, Chan-Ho Song, Min-Jae Hwang, Suhyeon Oh, Hyun-Wook Yoon, Jin-Seob Kim, Jae-Min Kim |
| 2022 | Interspeech | Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation. | Ryo Terashima, Ryuichi Yamamoto, Eunwoo Song, Yuma Shirahata, Hyun-Wook Yoon, Jae-Min Kim, Kentaro Tachibana |
| 2022 | Interspeech | Language Model-Based Emotion Prediction Methods for Emotional Speech Synthesis Systems. | Hyun-Wook Yoon, Ohsung Kwon, Hoyeon Lee, Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim, Min-Jae Hwang |
| 2021 | ICASSP | TTS-by-TTS: TTS-Driven Data Augmentation for Fast and High-Quality Speech Synthesis. | Min-Jae Hwang, Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim |
| 2021 | ICASSP | Parallel Waveform Synthesis Based on Generative Adversarial Networks with Voicing-Aware Conditional Discriminators. | Ryuichi Yamamoto, Eunwoo Song, Min-Jae Hwang, Jae-Min Kim |
| 2021 | Interspeech | Phrase Break Prediction with Bidirectional Encoder Representations in Japanese Text-to-Speech Synthesis. | Kosuke Futamata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
| 2021 | Interspeech | High-Fidelity Parallel WaveGAN with Multi-Band Harmonic-Plus-Noise Model. | Min-Jae Hwang, Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim |
| 2020 | ICASSP | Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit. | Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, Xu Tan |
| 2020 | ICASSP | Improving LPCNET-Based Text-to-Speech with Linear Prediction-Structured Mixture Density Network. | Min-Jae Hwang, Eunwoo Song, Ryuichi Yamamoto, Frank K. Soong, Hong-Goo Kang |
| 2020 | ICASSP | Semi-Supervised Speaker Adaptation for End-to-End Speech Synthesis with Pretrained Models. | Katsuki Inoue, Sunao Hara, Masanobu Abe, Tomoki Hayashi, Ryuichi Yamamoto, Shinji Watanabe |
| 2020 | ICASSP | Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram. | Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim |
| 2020 | Interspeech | Neural Text-to-Speech with a Modeling-by-Generation Excitation Vocoder. | Eunwoo Song, Min-Jae Hwang, Ryuichi Yamamoto, Jin-Seob Kim, Ohsung Kwon, Jae-Min Kim |
| 2019 | ASRU | A Comparative Study on Transformer vs RNN in Speech Applications. | Shigeki Karita, Xiaofei Wang, Shinji Watanabe, Takenori Yoshimura, Wangyou Zhang, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto |
| 2019 | Interspeech | Probability Density Distillation with Generative Adversarial Networks for High-Quality Parallel Waveform Generation. | Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim |
| 2013 | ICASSP | Robust on-line algorithm for real-time audio-to-score alignment based on a delayed decision and anticipation framework. | Ryuichi Yamamoto, Shinji Sako, Tadashi Kitamura |
| 2002 | CBMS | IFSIMS -Internet Frame ork Service for Intelligent Medical Systems. | Mitja Lenic, Peter Kokol, Ryuichi Yamamoto |
| 2002 | CBMS | A Framework for Dynamic Evidence Based Medicine using Data Mining. | Gou Masuda, Norihiro Sakamoto, Ryuichi Yamamoto |
| 2002 | CBMS | Mining Diabetes Database With Decision Trees and Association Rules. | Milan Zorman, Gou Masuda, Peter Kokol, Ryuichi Yamamoto, Bruno Stiglic |