ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts.
Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, Yong Jae Lee
Browse the full CVPR paper archive.
Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, Yong Jae Lee
Browse the full CVPR paper archive.