MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization.
Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma, Lu Yin, Qiao Xiao, Stavros Petridis, Shiwei Liu, Maja Pantic
Browse the full Interspeech paper archive.