GroundVLP: Harnessing Zero-Shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection.
Haozhan Shen, Tiancheng Zhao, Mingwei Zhu, Jianwei Yin
Browse the full AAAI paper archive.
Haozhan Shen, Tiancheng Zhao, Mingwei Zhu, Jianwei Yin
Browse the full AAAI paper archive.