STOA-VLP: Spatial-Temporal Modeling of Object and Action for Video-Language Pre-training.
Weihong Zhong, Mao Zheng, Duyu Tang, Xuan Luo, Heng Gong, Xiaocheng Feng, Bing Qin
Browse the full AAAI paper archive.
Weihong Zhong, Mao Zheng, Duyu Tang, Xuan Luo, Heng Gong, Xiaocheng Feng, Bing Qin
Browse the full AAAI paper archive.