Skip to content

SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization.

Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen, Chris Ngo, Truong-Son Hy

VenueA*EMNLP
Year2025
ProceedingsEMNLP

Browse the full EMNLP paper archive.