Spoken Moments: Learning Joint Audio-Visual Representations From Video Descriptions.
Mathew Monfort, SouYoung Jin, Alexander H. Liu, David Harwath, Rogrio Feris, James R. Glass, Aude Oliva
Browse the full CVPR paper archive.
Mathew Monfort, SouYoung Jin, Alexander H. Liu, David Harwath, Rogrio Feris, James R. Glass, Aude Oliva
Browse the full CVPR paper archive.