DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding.
Xiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng, Xiaofeng Wang, Yun Zheng, Xingang Wang
Browse the full ICCV paper archive.
Xiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng, Xiaofeng Wang, Yun Zheng, Xingang Wang
Browse the full ICCV paper archive.