MM-CARP: Multimodal Model with Cross-Modal Retrieval-Augmented and Visual Region Perception.
Junhao Guo, Chenhan Fu, Guoming Wang, Rongxing Lu, Dong Chen, Siliang Tang
Browse the full MMM paper archive.
Junhao Guo, Chenhan Fu, Guoming Wang, Rongxing Lu, Dong Chen, Siliang Tang
Browse the full MMM paper archive.