ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation.
Ali Athar, Xueqing Deng, Liang-Chieh Chen
Browse the full CVPR paper archive.
Ali Athar, Xueqing Deng, Liang-Chieh Chen
Browse the full CVPR paper archive.