How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Heads.
Ingeol Baek, Hwan Chang, Sunghyun Ryu, Hwanhee Lee
Browse the full EMNLP paper archive.
Ingeol Baek, Hwan Chang, Sunghyun Ryu, Hwanhee Lee
Browse the full EMNLP paper archive.