Skip to content

Textural or Textual: How Vision-Language Models Read Text in Images.

Hanzhang Wang, Qingyuan Ma

VenueA*ICML
Year2025
ProceedingsICML

Browse the full ICML paper archive.