Skip to content

MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens.

Jeong Hun Yeo, Hyeongseop Rha, Se Jin Park, Yong Man Ro

VenueA*ACL
Year2025
ProceedingsACL (Findings)

Browse the full ACL paper archive.