Nayana: A Foundation for Document-Centric Vision-Language Models via Multi-Task, Multimodal, and Multilingual Data Synthesis.
Adithya S. Kolavi, Samarth P, Vyoman Jain
Browse the full ICCV paper archive.
Adithya S. Kolavi, Samarth P, Vyoman Jain
Browse the full ICCV paper archive.