Skip to content

Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling.

Georgios Pantazopoulos, Malvina Nikandrou, Alessandro Suglia, Oliver Lemon, Arash Eshghi

VenueA*EMNLP
Year2024
ProceedingsEMNLP

Browse the full EMNLP paper archive.