Skip to content

AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning.

Jongsuk Kim, Jiwon Shin, Junmo Kim

Year2024
ProceedingsINTERSPEECH

Browse the full Interspeech paper archive.