Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman Mohamed
Browse the full ICLR paper archive.
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman Mohamed
Browse the full ICLR paper archive.