ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment.
Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee
Browse the full ACL paper archive.
Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee
Browse the full ACL paper archive.