Skip to content

Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding.

Yuanhao Xiong, Long Zhao, Boqing Gong, Ming-Hsuan Yang, Florian Schroff, Ting Liu, Cho-Jui Hsieh, Liangzhe Yuan

VenueA*ICLR
Year2024
ProceedingsICLR

Browse the full ICLR paper archive.