Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval.
Ruofan Hu, Yan Xia, Minjie Hong, Jieming Zhu, Bo Chen, Xiaoda Yang, Minghui Fang, Tao Jin
Browse the full Interspeech paper archive.
Ruofan Hu, Yan Xia, Minjie Hong, Jieming Zhu, Bo Chen, Xiaoda Yang, Minghui Fang, Tao Jin
Browse the full Interspeech paper archive.