Show and Speak: Directly Synthesize Spoken Description of Images.
Xinsheng Wang, Siyuan Feng, Jihua Zhu, Mark Hasegawa-Johnson, Odette Scharenborg
Browse the full ICASSP paper archive.
Xinsheng Wang, Siyuan Feng, Jihua Zhu, Mark Hasegawa-Johnson, Odette Scharenborg
Browse the full ICASSP paper archive.