Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes.
Alexandros Delitzas, Maria Parelli, Nikolas Hars, Georgios Vlassis, Sotirios-Konstantinos Anagnostidis, Gregor Bachmann, Thomas Hofmann
Browse the full BMVC paper archive.