ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learning.
Xiao Xu, Bei Li, Chenfei Wu, Shao-Yen Tseng, Anahita Bhiwandiwalla, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan
Browse the full ACL paper archive.