MIST : Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question Answering.
Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, Mike Zheng Shou
Browse the full CVPR paper archive.
Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, Mike Zheng Shou
Browse the full CVPR paper archive.