Skip to content

ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation.

Ali Athar, Xueqing Deng, Liang-Chieh Chen

VenueA*CVPR
Year2025
ProceedingsCVPR

Browse the full CVPR paper archive.