Segmental SpeechCLIP: Utilizing Pretrained Image-text Models for Audio-Visual Learning.
Saurabhchand Bhati, Jess Villalba, Laureano Moro-Velzquez, Thomas Thebaud, Najim Dehak
Browse the full Interspeech paper archive.
Saurabhchand Bhati, Jess Villalba, Laureano Moro-Velzquez, Thomas Thebaud, Najim Dehak
Browse the full Interspeech paper archive.