GTA: Towards Generative Text-To-Audio Retrieval via Multi-Scale Tokenizer.
Minghui Fang, Shengpeng Ji, Jialong Zuo, Xize Cheng, Wenrui Liu, Xiaoda Yang, Ruofan Hu, Jieming Zhu, Zhou Zhao
Browse the full Interspeech paper archive.
Minghui Fang, Shengpeng Ji, Jialong Zuo, Xize Cheng, Wenrui Liu, Xiaoda Yang, Ruofan Hu, Jieming Zhu, Zhou Zhao
Browse the full Interspeech paper archive.