Skip to content

JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse.

Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma, Yitao Liang

VenueA*ACL
Year2025
ProceedingsACL (Findings)

Browse the full ACL paper archive.