Skip to content

From Pixels to Voice: A Simple and Efficient End-to-End Spoken Image Description Approach via Vision Codec Language Models.

Chung Tran, Sakriani Sakti

Year2025
ProceedingsICASSP

Browse the full ICASSP paper archive.