Vx2Text: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs.
Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, Lorenzo Torresani
Browse the full CVPR paper archive.
Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, Lorenzo Torresani
Browse the full CVPR paper archive.