Skip to content

Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman Mohamed

VenueA*ICLR
Year2022
ProceedingsICLR

Browse the full ICLR paper archive.