Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model.
Jinlong Xue, Yayue Deng, Yicheng Han, Yingming Gao, Ya Li
Browse the full Interspeech paper archive.
Jinlong Xue, Yayue Deng, Yicheng Han, Yingming Gao, Ya Li
Browse the full Interspeech paper archive.