Learning CLIP Guided Visual-Text Fusion Transformer for Video-based Pedestrian Attribute Recognition.
Jun Zhu, Jiandong Jin, Zihan Yang, Xiaohao Wu, Xiao Wang
Browse the full CVPR paper archive.
Jun Zhu, Jiandong Jin, Zihan Yang, Xiaohao Wu, Xiao Wang
Browse the full CVPR paper archive.