| 2026 | ACL | Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions. | Dongwook Lee, Eunwoo Song, Che Hyun Lee, Heeseung Kim, Sungroh Yoon |
| 2025 | Interspeech | RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching. | Hyun Joon Park, Jeongmin Liu, Jin Sob Kim, Jeong Yeol Yang, Sung Won Han, Eunwoo Song |
| 2024 | ICASSP | Enhancing Multilingual TTS with Voice Conversion Based Data Augmentation and Posterior Embedding. | Hyun-Wook Yoon, Jin-Seob Kim, Ryuichi Yamamoto, Ryo Terashima, Chan-Ho Song, Jae-Min Kim, Eunwoo Song |
| 2023 | ICASSP | Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech Synthesis. | Yuma Shirahata, Ryuichi Yamamoto, Eunwoo Song, Ryo Terashima, Jae-Min Kim, Kentaro Tachibana |
| 2023 | Interspeech | Pruning Self-Attention for Zero-Shot Multi-Speaker Text-to-Speech. | Hyungchan Yoon, Changhwan Kim, Eunwoo Song, Hyun-Wook Yoon, Hong-Goo Kang |
| 2022 | Interspeech | TTS-by-TTS 2: Data-Selective Augmentation for Neural Speech Synthesis Using Ranking Support Vector Machine with Variational Autoencoder. | Eunwoo Song, Ryuichi Yamamoto, Ohsung Kwon, Chan-Ho Song, Min-Jae Hwang, Suhyeon Oh, Hyun-Wook Yoon, Jin-Seob Kim, Jae-Min Kim |
| 2022 | Interspeech | Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation. | Ryo Terashima, Ryuichi Yamamoto, Eunwoo Song, Yuma Shirahata, Hyun-Wook Yoon, Jae-Min Kim, Kentaro Tachibana |
| 2022 | Interspeech | Language Model-Based Emotion Prediction Methods for Emotional Speech Synthesis Systems. | Hyun-Wook Yoon, Ohsung Kwon, Hoyeon Lee, Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim, Min-Jae Hwang |
| 2021 | ICASSP | TTS-by-TTS: TTS-Driven Data Augmentation for Fast and High-Quality Speech Synthesis. | Min-Jae Hwang, Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim |
| 2021 | ICASSP | Parallel Waveform Synthesis Based on Generative Adversarial Networks with Voicing-Aware Conditional Discriminators. | Ryuichi Yamamoto, Eunwoo Song, Min-Jae Hwang, Jae-Min Kim |
| 2021 | Interspeech | High-Fidelity Parallel WaveGAN with Multi-Band Harmonic-Plus-Noise Model. | Min-Jae Hwang, Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim |
| 2021 | Interspeech | LiteTTS: A Lightweight Mel-Spectrogram-Free Text-to-Wave Synthesizer Based on Generative Adversarial Networks. | Huu-Kim Nguyen, Kihyuk Jeong, Seyun Um, Min-Jae Hwang, Eunwoo Song, Hong-Goo Kang |
| 2020 | ICASSP | Improving LPCNET-Based Text-to-Speech with Linear Prediction-Structured Mixture Density Network. | Min-Jae Hwang, Eunwoo Song, Ryuichi Yamamoto, Frank K. Soong, Hong-Goo Kang |
| 2020 | ICASSP | Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram. | Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim |
| 2020 | Interspeech | Neural Text-to-Speech with a Modeling-by-Generation Excitation Vocoder. | Eunwoo Song, Min-Jae Hwang, Ryuichi Yamamoto, Jin-Seob Kim, Ohsung Kwon, Jae-Min Kim |
| 2020 | MMSP | Speaker-Adaptive Neural Vocoders for Parametric Speech Synthesis Systems. | Eunwoo Song, Jin-Seob Kim, Kyungguen Byun, Hong-Goo Kang |
| 2019 | Interspeech | Probability Density Distillation with Generative Adversarial Networks for High-Quality Parallel Waveform Generation. | Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim |
| 2018 | ICASSP | Modeling-By-Generation-Structured Noise Compensation Algorithm for Glottal Vocoding Speech Synthesis System. | Min-Jae Hwang, Eunwoo Song, Kyungguen Byun, Hong-Goo Kang |
| 2018 | Interspeech | A Unified Framework for the Generation of Glottal Signals in Deep Learning-based Parametric Speech Synthesis Systems. | Min-Jae Hwang, Eunwoo Song, Jin-Seob Kim, Hong-Goo Kang |
| 2018 | Interspeech | Acoustic Modeling Using Adversarially Trained Variational Recurrent Neural Network for Speech Synthesis. | Joun Yeop Lee, Sung Jun Cheon, Byoung Jin Choi, Nam Soo Kim, Eunwoo Song |
| 2017 | ASRU | Perceptual quality and modeling accuracy of excitation parameters in DLSTM-based speech synthesis systems. | Eunwoo Song, Frank K. Soong, Hong-Goo Kang |
| 2016 | Interspeech | Improved Time-Frequency Trajectory Excitation Vocoder for DNN-Based Speech Synthesis. | Eunwoo Song, Frank K. Soong, Hong-Goo Kang |
| 2015 | ICASSP | Improved time-frequency trajectory excitation modeling for a statistical parametric speech synthesis system. | Eunwoo Song, Young-Sun Joo, Hong-Goo Kang |
| 2015 | Interspeech | Deep neural network-based statistical parametric speech synthesis system using improved time-frequency trajectory excitation model. | Eunwoo Song, Hong-Goo Kang |