Skip to content

CLIP4VideoCap: Rethinking Clip for Video Captioning with Multiscale Temporal Fusion and Commonsense Knowledge.

Tanvir Mahmud, Feng Liang, Yaling Qing, Diana Marculescu

Year2023
ProceedingsICASSP

Browse the full ICASSP paper archive.