Skip to content

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval.

Ruofan Hu, Yan Xia, Minjie Hong, Jieming Zhu, Bo Chen, Xiaoda Yang, Minghui Fang, Tao Jin

Year2025
ProceedingsINTERSPEECH

Browse the full Interspeech paper archive.