Skip to content

What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions.

Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas, Shih-Fu Chang, Rogrio Feris, James R. Glass, Hilde Kuehne

VenueA*CVPR
Year2024
ProceedingsCVPR

Browse the full CVPR paper archive.