LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.
Shaolei Zhang, Qingkai Fang, Zhe Yang, Yang Feng
Browse the full ICLR paper archive.
Shaolei Zhang, Qingkai Fang, Zhe Yang, Yang Feng
Browse the full ICLR paper archive.