Skip to content

VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models.

Harshit, Tolga Tasdizen

VenueAWACV
Year2025
ProceedingsWACV (Workshops)

Browse the full WACV paper archive.