Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding.
Seil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae Hwang
Browse the full CVPR paper archive.
Seil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae Hwang
Browse the full CVPR paper archive.