Skip to content

Andrew Rouditchenko

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

17

Venues

8

Active years

2018–2025

Best venue rank

A*

Where they publish

Papers

17 indexed papers, newest first.

YearVenueTitleAuthors
2025ASRUOmni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?Andrew Rouditchenko, Saurabhchand Bhati, Edson Araujo, Samuel Thomas, Hilde Kuehne, Rogrio Feris, James R. Glass
2025CVPRCAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment.Edson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati, Samuel Thomas, Brian Kingsbury, Leonid Karlinsky, Rogrio Feris, James R. Glass, Hilde Kuehne
2024CVPRWhat, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions.Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas, Shih-Fu Chang, Rogrio Feris, James R. Glass, Hilde Kuehne
2024ECCVAV-CPL: Continuous Pseudo-labeling for Audio-Visual Speech Recognition.Andrew Rouditchenko, Ronan Collobert, Tatiana Likhomanenko
2024InterspeechWhisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation.Andrew Rouditchenko, Yuan Gong, Samuel Thomas, Leonid Karlinsky, Hilde Kuehne, Rogrio Feris, James Glass
2023ICASSPC2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval.Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas, Rogrio Feris, Brian Kingsbury, Leonid Karlinsky, David Harwath, Hilde Kuehne, James R. Glass
2023ICLRContrastive Audio-Visual Masked Autoencoder.Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, James R. Glass
2023InterspeechComparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages.Andrew Rouditchenko, Sameer Khurana, Samuel Thomas, Rogrio Feris, Leonid Karlinsky, Hilde Kuehne, David Harwath, Brian Kingsbury, James R. Glass
2022ACLCross-Modal Discrete Representation Learning.Alexander H. Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, James R. Glass
2022CVPREverything at Once - Multi-modal Fusion Transformer for Video Retrieval.Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Hilde Kuehne
2021ICCVMultimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos.Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang
2021InterspeechSpoken ObjectNet: A Bias-Controlled Spoken Caption Dataset.Ian Palmer, Andrew Rouditchenko, Andrei Barbu, Boris Katz, James R. Glass
2021InterspeechCascaded Multilingual Audio-Visual Learning from Videos.Andrew Rouditchenko, Angie W. Boggust, David Harwath, Samuel Thomas, Hilde Kuehne, Brian Chen, Rameswar Panda, Rogrio Feris, Brian Kingsbury, Michael Picheny, James R. Glass
2021InterspeechAVLnet: Learning Audio-Visual Language Representations from Instructional Videos.Andrew Rouditchenko, Angie W. Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogrio Schmidt Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, James R. Glass
2019CVPRSelf-Supervised Segmentation and Source Separation on Videos.Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh H. McDermott, Antonio Torralba
2019ICASSPSelf-supervised Audio-visual Co-segmentation.Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh H. McDermott, Antonio Torralba
2018ECCVThe Sound of Pixels.Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh H. McDermott, Antonio Torralba