CLIP-Guided Vision-Language Pre-training for Question Answering in 3D Scenes.
Maria Parelli, Alexandros Delitzas, Nikolas Hars, Georgios Vlassis, Sotiris Anagnostidis, Gregor Bachmann, Thomas Hofmann
Browse the full CVPR paper archive.
Maria Parelli, Alexandros Delitzas, Nikolas Hars, Georgios Vlassis, Sotiris Anagnostidis, Gregor Bachmann, Thomas Hofmann
Browse the full CVPR paper archive.