| 2025 | ASRU | CAVIARES: Corpus for Audio-Visual Expressive Voice Agent. | Jinsheng Chen, Yuki Saito, Dong Yang, Naoko Tanji, Hironori Doi, Byeongseon Park, Yuma Shirahata, Kentaro Tachibana, Hiroshi Saruwatari |
| 2025 | ICASSP | Description-Based Controllable Text-to-Speech With Cross-Lingual Voice Control. | Ryuichi Yamamoto, Yuma Shirahata, Masaya Kawamura, Kentaro Tachibana |
| 2025 | Interspeech | BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing. | Masaya Kawamura, Takuya Hasumi, Yuma Shirahata, Ryuichi Yamamoto |
| 2025 | Interspeech | Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning. | Hien Ohnaka, Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto |
| 2025 | Interspeech | SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch. | Ryo Terashima, Yuma Shirahata, Masaya Kawamura |
| 2024 | ICASSP | PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language Descriptions. | Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, Kentaro Tachibana |
| 2024 | Interspeech | LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning. | Masaya Kawamura, Ryuichi Yamamoto, Yuma Shirahata, Takuya Hasumi, Kentaro Tachibana |
| 2024 | Interspeech | Universal Score-based Speech Enhancement with High Content Preservation. | Robin Scheibler, Yusuke Fujita, Yuma Shirahata, Tatsuya Komatsu |
| 2024 | Interspeech | Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data. | Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
| 2023 | ICASSP | Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform. | Masaya Kawamura, Yuma Shirahata, Ryuichi Yamamoto, Kentaro Tachibana |
| 2023 | ICASSP | Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech Synthesis. | Yuma Shirahata, Ryuichi Yamamoto, Eunwoo Song, Ryo Terashima, Jae-Min Kim, Kentaro Tachibana |
| 2022 | Interspeech | Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation. | Ryo Terashima, Ryuichi Yamamoto, Eunwoo Song, Yuma Shirahata, Hyun-Wook Yoon, Jae-Min Kim, Kentaro Tachibana |
| 2020 | Interspeech | Discriminative Method to Extract Coarse Prosodic Structure and its Application for Statistical Phrase/Accent Command Estimation. | Yuma Shirahata, Daisuke Saito, Nobuaki Minematsu |