MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.
Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, Filip Ilievski
Browse the full ICLR paper archive.
Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, Filip Ilievski
Browse the full ICLR paper archive.