Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling.
Georgios Pantazopoulos, Malvina Nikandrou, Alessandro Suglia, Oliver Lemon, Arash Eshghi
Browse the full EMNLP paper archive.
Georgios Pantazopoulos, Malvina Nikandrou, Alessandro Suglia, Oliver Lemon, Arash Eshghi
Browse the full EMNLP paper archive.