| 2025 | AAAI | Scaling Diffusion Mamba with Bidirectional SSMs for Efficient 3D Shape Generation. | Shentong Mo |
| 2025 | AAAI | The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning. | Shentong Mo |
| 2025 | CVPR | Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows. | Shentong Mo, Yibing Song |
| 2025 | ICASSP | DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap. | Shentong Mo, Zehua Chen, Fan Bao, Jun Zhu |
| 2025 | ICLR | pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation. | Shentong Mo, Xufang Luo, Dongsheng Li |
| 2025 | ICML | GMAIL: Generative Modality Alignment for generated Image Learning. | Shentong Mo, Sukmin Yun |
| 2024 | CVPR | MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers. | Tanvir Mahmud, Shentong Mo, Yapeng Tian, Diana Marculescu |
| 2024 | CVPR | Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions Through Masked Modeling. | Shentong Mo, Pedro Morgado |
| 2024 | ECCV | Audio-Visual Generalized Zero-Shot Learning the Easy Way. | Shentong Mo, Pedro Morgado |
| 2024 | ECCV | Fast Training of Diffusion Transformer with Extreme Masking for 3D Point Clouds Generation. | Shentong Mo, Enze Xie, Yue Wu, Junsong Chen, Matthias Niener, Zhenguo Li |
| 2024 | ECCV | DailyMAE: Towards Pretraining Masked Autoencoders in One Day. | Jiantao Wu, Shentong Mo, Sara Atito, Zhenhua Feng, Josef Kittler, Muhammad Awais |
| 2024 | ECCV | Audio-Synchronized Visual Animation. | Lin Zhang, Shentong Mo, Yijing Zhang, Pedro Morgado |
| 2024 | ICASSP | Tree of Uncertain Thoughts Reasoning for Large Language Models. | Shentong Mo, Miao Xin |
| 2024 | ICIP | Masked Momentum Contrastive Learning for Semantic Understanding by Observation. | Jiantao Wu, Shentong Mo, Sara Atito, Zhenhua Feng, Josef Kittler, Syed Sameed Husain, Muhammad Awais |
| 2023 | BMVC | Variational Autoencoders with Decremental Information Bottleneck for Disentanglement. | Jiantao Wu, Shentong Mo, Xingshen Zhang, Muhammad Awais, Sara Ahmed, Zhenhua Feng, Lin Wang, Xiang Yang |
| 2023 | CVPR | Audio-Visual Grouping Network for Sound Localization from Mixtures. | Shentong Mo, Yapeng Tian |
| 2023 | ICCV | Class-Incremental Grouping Network for Continual Audio-Visual Learning. | Shentong Mo, Weiguo Pian, Yapeng Tian |
| 2023 | ICCV | Audio-Visual Class-Incremental Learning. | Weiguo Pian, Shentong Mo, Yunhui Guo, Yapeng Tian |
| 2023 | ICML | A Unified Audio-Visual Learning Framework for Localization, Separation, and Recognition. | Shentong Mo, Pedro Morgado |
| 2023 | WACV | Representation Disentanglement in Generative Models with Contrastive Learning. | Shentong Mo, Zhun Sun, Chao Li |
| 2023 | WACV | Multi-level Contrastive Learning for Self-Supervised Vision Transformers. | Shentong Mo, Zhun Sun, Chao Li |
| 2022 | BMVC | Rethinking Prototypical Contrastive Learning through Alignment, Uniformity and Correlation. | Shentong Mo, Zhun Sun, Chao Li |
| 2022 | ECCV | Unitail: Detecting, Reading, and Matching in Retail Scene. | Fangyi Chen, Han Zhang, Zaiwang Li, Jiachen Dou, Shentong Mo, Hao Chen, Yongxin Zhang, Uzair Ahmed, Chenchen Zhu, Marios Savvides |
| 2022 | ECCV | Localizing Visual Sounds the Easy Way. | Shentong Mo, Pedro Morgado |
| 2021 | BMVC | Siamese Prototypical Contrastive Learning. | Shentong Mo, Zhun Sun, Chao Li |
| 2021 | BMVC | Point3D: tracking actions as moving points with 3D CNNs. | Shentong Mo, Jingfei Xia, Xiaoqing Tan, Bhiksha Raj |
| 2021 | CVPR | Long-Term Head Pose Forecasting Conditioned on the Gaze-Guiding Prior. | Shentong Mo, Miao Xin |
| 2021 | CVPR | EVA-GCN: Head Pose Estimation Based on Graph Convolutional Networks. | Miao Xin, Shentong Mo, Yuanze Lin |