Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data.
Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana
Browse the full Interspeech paper archive.
Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana
Browse the full Interspeech paper archive.