Skip to content

Improving Audio Captioning Models with Fine-Grained Audio Features, Text Embedding Supervision, and LLM Mix-Up Augmentation.

Shih-Lun Wu, Xuankai Chang, Gordon Wichern, Jee-Weon Jung, Franois G. Germain, Jonathan Le Roux, Shinji Watanabe

Year2024
ProceedingsICASSP

Browse the full ICASSP paper archive.