MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens.
Jeong Hun Yeo, Hyeongseop Rha, Se Jin Park, Yong Man Ro
Browse the full ACL paper archive.
Jeong Hun Yeo, Hyeongseop Rha, Se Jin Park, Yong Man Ro
Browse the full ACL paper archive.