| 2024 | ACL | Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition. | Yukiya Hono, Koh Mitsuda, Tianyu Zhao, Kentaro Mitsui, Toshiaki Wakatsuki, Kei Sawada |
| 2024 | COLING | Release of Pre-Trained Models for the Japanese Language. | Kei Sawada, Tianyu Zhao, Makoto Shing, Kentaro Mitsui, Akio Kaga, Yukiya Hono, Toshiaki Wakatsuki, Koh Mitsuda |
| 2024 | EMNLP | PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems. | Kentaro Mitsui, Koh Mitsuda, Toshiaki Wakatsuki, Yukiya Hono, Kei Sawada |
| 2023 | Interspeech | UniFLG: Unified Facial Landmark Generator from Text or Speech. | Kentaro Mitsui, Yukiya Hono, Kei Sawada |
| 2022 | Interspeech | MSR-NV: Neural Vocoder Using Multiple Sampling Rates. | Kentaro Mitsui, Kei Sawada |
| 2022 | Interspeech | End-to-End Text-to-Speech Based on Latent Representation of Speaking Styles Using Spontaneous Dialogue. | Kentaro Mitsui, Tianyu Zhao, Kei Sawada, Yukiya Hono, Yoshihiko Nankaku, Keiichi Tokuda |
| 2020 | Interspeech | Multi-Speaker Text-to-Speech Synthesis Using Deep Gaussian Processes. | Kentaro Mitsui, Tomoki Koriyama, Hiroshi Saruwatari |