Skip to content

Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language.

Mark Hamilton, Andrew Zisserman, John R. Hershey, William T. Freeman

VenueA*CVPR
Year2024
ProceedingsCVPR

Browse the full CVPR paper archive.