Skip to content

CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion.

Shoubin Yu, Jaehong Yoon, Mohit Bansal

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.