Enrich and Detect: Video Temporal Grounding With Multimodal Llms.
Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa, Lorenzo Torresani, Triantafyllos Afouras
Browse the full ICCV paper archive.
Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa, Lorenzo Torresani, Triantafyllos Afouras
Browse the full ICCV paper archive.