Skip to content

MAPS: Joint Multimodal Attention and POS Sequence Generation for Video Captioning.

Cong Zou, Xuchen Wang, Yaosi Hu, Zhenzhong Chen, Shan Liu

VenueCVCIP
Year2021
ProceedingsVCIP

Browse the full VCIP paper archive.