Unsupervised Audio-Caption Aligning Learns Correspondences Between Individual Sound Events and Textual Phrases.
Huang Xie, Okko Rsnen, Konstantinos Drossos, Tuomas Virtanen
Browse the full ICASSP paper archive.
Huang Xie, Okko Rsnen, Konstantinos Drossos, Tuomas Virtanen
Browse the full ICASSP paper archive.