Skip to content

Segmental SpeechCLIP: Utilizing Pretrained Image-text Models for Audio-Visual Learning.

Saurabhchand Bhati, Jess Villalba, Laureano Moro-Velzquez, Thomas Thebaud, Najim Dehak

Year2023
ProceedingsINTERSPEECH

Browse the full Interspeech paper archive.