JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse.
Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma, Yitao Liang
Browse the full ACL paper archive.
Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma, Yitao Liang
Browse the full ACL paper archive.