Transformer-Based Video Front-Ends for Audio-Visual Speech Recognition for Single and Muti-Person Video.
Dmitriy Serdyuk, Otavio Braga, Olivier Siohan
Browse the full Interspeech paper archive.
Dmitriy Serdyuk, Otavio Braga, Olivier Siohan
Browse the full Interspeech paper archive.