Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding.
Zhanpeng Chen, Mingxiao Li, Ziyang Chen, Nan Du, Xiaolong Li, Yuexian Zou
Browse the full ACL paper archive.
Zhanpeng Chen, Mingxiao Li, Ziyang Chen, Nan Du, Xiaolong Li, Yuexian Zou
Browse the full ACL paper archive.