Skip to content
cs-conference-ranking
.org
By subfield
By rank
Methodology
⌕
Search 971 venues
Home
/
Interspeech
/
Paper
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning.
Jongsuk Kim
,
Jiwon Shin
,
Junmo Kim
Venue
A
Interspeech
Year
2024
Proceedings
INTERSPEECH
DBLP record
conf/interspeech/KimS024 ↗
Browse the full
Interspeech paper archive
.