Skip to content

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding.

Xiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng, Xiaofeng Wang, Yun Zheng, Xingang Wang

VenueA*ICCV
Year2025
ProceedingsICCV

Browse the full ICCV paper archive.