Skip to content

OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer.

Lu Zhang, Tiancheng Zhao, Heting Ying, Yibo Ma, Kyusong Lee

VenueA*EMNLP
Year2024
ProceedingsEMNLP

Browse the full EMNLP paper archive.