Skip to content

Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Hang Zhang, Xin Li, Lidong Bing

VenueA*EMNLP
Year2023
ProceedingsEMNLP (Demos)

Browse the full EMNLP paper archive.