PERCEIVER-VL: Efficient Vision-and-Language Modeling with Iterative Latent Attention.
Zineng Tang, Jaemin Cho, Jie Lei, Mohit Bansal
Browse the full WACV paper archive.
Zineng Tang, Jaemin Cho, Jie Lei, Mohit Bansal
Browse the full WACV paper archive.