Skip to content

Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data.

Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana

Year2024
ProceedingsINTERSPEECH

Browse the full Interspeech paper archive.