| 2025 | ASRU | CAVIARES: Corpus for Audio-Visual Expressive Voice Agent. | Jinsheng Chen, Yuki Saito, Dong Yang, Naoko Tanji, Hironori Doi, Byeongseon Park, Yuma Shirahata, Kentaro Tachibana, Hiroshi Saruwatari |
| 2025 | ICASSP | Description-Based Controllable Text-to-Speech With Cross-Lingual Voice Control. | Ryuichi Yamamoto, Yuma Shirahata, Masaya Kawamura, Kentaro Tachibana |
| 2024 | ICASSP | PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language Descriptions. | Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, Kentaro Tachibana |
| 2024 | Interspeech | Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment. | Takuto Igarashi, Yuki Saito, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari |
| 2024 | Interspeech | LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning. | Masaya Kawamura, Ryuichi Yamamoto, Yuma Shirahata, Takuya Hasumi, Kentaro Tachibana |
| 2024 | Interspeech | SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark. | Yuki Saito, Takuto Igarashi, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari |
| 2024 | Interspeech | Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data. | Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
| 2023 | ICASSP | Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform. | Masaya Kawamura, Yuma Shirahata, Ryuichi Yamamoto, Kentaro Tachibana |
| 2023 | ICASSP | Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech Synthesis. | Yuma Shirahata, Ryuichi Yamamoto, Eunwoo Song, Ryo Terashima, Jae-Min Kim, Kentaro Tachibana |
| 2023 | ICASSP | Nonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs. | Reo Yoneyama, Ryuichi Yamamoto, Kentaro Tachibana |
| 2023 | Interspeech | CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center. | Yuki Saito, Eiji Iimori, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari |
| 2023 | Interspeech | ChatGPT-EDSS: Empathetic Dialogue Speech Synthesis Trained from ChatGPT-derived Context Word Embeddings. | Yuki Saito, Shinnosuke Takamichi, Eiji Iimori, Kentaro Tachibana, Hiroshi Saruwatari |
| 2022 | Interspeech | Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History. | Yuto Nishimura, Yuki Saito, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari |
| 2022 | Interspeech | A Unified Accent Estimation Method Based on Multi-Task Learning for Japanese Text-to-Speech. | Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
| 2022 | Interspeech | DRSpeech: Degradation-Robust Text-to-Speech Synthesis with Frame-Level and Utterance-Level Acoustic Representation Learning. | Takaaki Saeki, Kentaro Tachibana, Ryuichi Yamamoto |
| 2022 | Interspeech | STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent. | Yuki Saito, Yuto Nishimura, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari |
| 2022 | Interspeech | Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation. | Ryo Terashima, Ryuichi Yamamoto, Eunwoo Song, Yuma Shirahata, Hyun-Wook Yoon, Jae-Min Kim, Kentaro Tachibana |
| 2021 | Interspeech | Phrase Break Prediction with Bidirectional Encoder Representations in Japanese Text-to-Speech Synthesis. | Kosuke Futamata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
| 2020 | Interspeech | Face2Speech: Towards Multi-Speaker Text-to-Speech Synthesis Using an Embedding Vector Predicted from a Face Image. | Shunsuke Goto, Kotaro Onishi, Yuki Saito, Kentaro Tachibana, Koichiro Mori |
| 2018 | ECCV | Full-Body High-Resolution Anime Generation with Progressive Structure-Conditional Generative Adversarial Networks. | Koichi Hamada, Kentaro Tachibana, Tianqi Li, Hiroto Honda, Yusuke Uchida |
| 2018 | ICASSP | An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic Features. | Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2018 | ICASSP | An Investigation of Noise Shaping with Perceptual Weighting for Wavenet-Based Speech Generation. | Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2017 | ASRU | Subband wavenet with overlapped single-sideband filterbanks. | Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2016 | Interspeech | Model Integration for HMM- and DNN-Based Speech Synthesis Using Product-of-Experts Framework. | Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2009 | ICASSP | Source adaptive blind signal extraction using closed-form ICA for hands-free robot spoken dialogue system. | Yu Takahashi, Hiroshi Saruwatari, Yuki Fujihara, Kentaro Tachibana, Yoshimitsu Mori, Shigeki Miyabe, Kiyohiro Shikano, Akira Tanaka |
| 2007 | ICASSP | Efficient Blind Source Separation Combining Closed-Form Second-Order ICA and Nonclosed-Form Higher-Order ICA. | Kentaro Tachibana, Hiroshi Saruwatari, Yoshimitsu Mori, Shigeki Miyabe, Kiyohiro Shikano, Akira Tanaka |