Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language.
Mark Hamilton, Andrew Zisserman, John R. Hershey, William T. Freeman
Browse the full CVPR paper archive.
Mark Hamilton, Andrew Zisserman, John R. Hershey, William T. Freeman
Browse the full CVPR paper archive.