Skip to content

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos.

Qixin Wang, Songtao Zhou, Zeyu Jin, Chenglin Guo, Shikun Sun, Xiaoyu Qin

VenueBIJCNN
Year2025
ProceedingsIJCNN

Browse the full IJCNN paper archive.