Skip to content

Spoken Moments: Learning Joint Audio-Visual Representations From Video Descriptions.

Mathew Monfort, SouYoung Jin, Alexander H. Liu, David Harwath, Rogrio Feris, James R. Glass, Aude Oliva

VenueA*CVPR
Year2021
ProceedingsCVPR

Browse the full CVPR paper archive.