Skip to content

ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts.

Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, Yong Jae Lee

VenueA*CVPR
Year2024
ProceedingsCVPR

Browse the full CVPR paper archive.