Skip to content

TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs.

Zijian Zhang, Xuhui Zheng, Xuecheng Wu, Chong Peng, Xuezhi Cao

VenueA*CVPR
Year2025
ProceedingsCVPR Workshops

Browse the full CVPR paper archive.