| 2026 | ACL | Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations. | Deok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, Seong-Whan Lee |
| 2026 | ACL | ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment. | Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee |
| 2025 | EMNLP | FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Control. | Seung-Bin Kim, Junhyeok Cha, Hyung-Seok Oh, Heejin Choi, Seong-Whan Lee |
| 2025 | ICASSP | JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis. | Junhyeok Cha, Seung-Bin Kim, Hyung-Seok Oh, Seong-Whan Lee |
| 2025 | ICASSP | FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching. | Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee |
| 2025 | Interspeech | DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech. | Deok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, Seong-Whan Lee |
| 2025 | Interspeech | EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification. | Deok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, Seong-Whan Lee |
| 2025 | Interspeech | Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech. | Nam-Gyu Kim, Deok-Hyeon Cho, Seung-Bin Kim, Seong-Whan Lee |
| 2024 | ICASSP | TranSentence: speech-to-speech Translation via Language-Agnostic Sentence-Level Speech Encoding without Language-Parallel Data. | Seung-Bin Kim, Sang-Hoon Lee, Seong-Whan Lee |
| 2024 | Interspeech | EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech. | Deok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, Sang-Hoon Lee, Seong-Whan Lee |
| 2024 | SMC | PromotiCon: Prompt-based Emotion Controllable Text-to-Speech via Prompt Generation and Matching. | Ji-Eun Lee, Seung-Bin Kim, Deok-Hyeon Cho, Seong-Whan Lee |
| 2022 | ICASSP | EMOQ-TTS: Emotion Intensity Quantization for Fine-Grained Controllable Emotional Text-to-Speech. | Chae-Bin Im, Sang-Hoon Lee, Seung-Bin Kim, Seong-Whan Lee |