| 2025 | ACL | MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities. | Savya Khosla, Aditi Tiwari, Kushal Kafle, Simon Jenni, Handong Zhao, John P. Collomosse, Jing Shi |
| 2025 | CVPR | The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers. | Daiqing Qi, Handong Zhao, Jing Shi, Simon Jenni, Yifei Fan, Franck Dernoncourt, Scott Cohen, Sheng Li |
| 2025 | CVPR | ViDROP: Video Dense Representation through Spatio-Temporal Sparsity. | Sepehr Sameni, Simon Jenni, Paolo Favaro |
| 2025 | ICCV | Improving Large Vision and Language Models by Learning from a Panel of Peers. | Jefferson Hernandez, Jing Shi, Simon Jenni, Vicente Ordonez, Kushal Kafle |
| 2024 | AAAI | No More Shortcuts: Realizing the Potential of Temporal Self-Supervision. | Ishan Rajendrakumar Dave, Simon Jenni, Mubarak Shah |
| 2024 | CVPR | Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models. | Gihyun Kwon, Simon Jenni, Dingzeyu Li, Joon-Young Lee, Jong Chul Ye, Fabian Caba Heilbron |
| 2024 | CVPR | Building Vision-Language Models on Solid Foundations with Masked Distillation. | Sepehr Sameni, Kushal Kafle, Hao Tan, Simon Jenni |
| 2024 | ECCV | Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets. | Ishan Rajendrakumar Dave, Fabian Caba Heilbron, Mubarak Shah, Simon Jenni |
| 2024 | ECCV | FineMatch: Aspect-Based Fine-Grained Image and Text Mismatch Detection and Correction. | Hang Hua, Jing Shi, Kushal Kafle, Simon Jenni, Daoan Zhang, John P. Collomosse, Scott Cohen, Jiebo Luo |
| 2023 | AAAI | Audio-Visual Contrastive Learning with Temporal Self-Supervision. | Simon Jenni, Alexander Black, John P. Collomosse |
| 2023 | AAAI | Representation Learning by Detecting Incorrect Location Embeddings. | Sepehr Sameni, Simon Jenni, Paolo Favaro |
| 2023 | CVPR | EKILA: Synthetic Media Provenance and Attribution for Generative Art. | Kar Balan, Shruti Agarwal, Simon Jenni, Andy Parsons, Andrew Gilbert, John P. Collomosse |
| 2023 | CVPR | Meta-Personalizing Vision-Language Models to Find Named Instances in Video. | Chun-Hsiao Yeh, Bryan C. Russell, Josef Sivic, Fabian Caba Heilbron, Simon Jenni |
| 2023 | ICCV | VADER: Video Alignment Differencing and Retrieval. | Alexander Black, Simon Jenni, Tu Bui, Md. Mehrab Tanjim, Stefano Petrangeli, Ritwik Sinha, Viswanathan Swaminathan, John P. Collomosse |
| 2023 | ICCV | Spatio-Temporal Crop Aggregation for Video Representation Learning. | Sepehr Sameni, Simon Jenni, Paolo Favaro |
| 2021 | BMVC | Learning to Deblur and Rotate Motion-Blurred Faces. | Givi Meishvili, Attila Szab, Simon Jenni, Paolo Favaro |
| 2021 | ICCV | Time-Equivariant Contrastive Video Representation Learning. | Simon Jenni, Hailin Jin |
| 2020 | ACCV | Self-supervised Multi-view Synchronization Learning for 3D Pose Estimation. | Simon Jenni, Paolo Favaro |
| 2020 | CVPR | Steering Self-Supervised Feature Learning Beyond Local Pixel Statistics. | Simon Jenni, Hailin Jin, Paolo Favaro |
| 2020 | CVPR | Learning to Have an Ear for Face Super-Resolution. | Givi Meishvili, Simon Jenni, Paolo Favaro |
| 2020 | ECCV | Video Representation Learning by Recognizing Temporal Transformations. | Simon Jenni, Givi Meishvili, Paolo Favaro |
| 2019 | CVPR | On Stabilizing Generative Adversarial Training With Noise. | Simon Jenni, Paolo Favaro |
| 2018 | CVPR | Self-Supervised Feature Learning by Learning to Spot Artifacts. | Simon Jenni, Paolo Favaro |
| 2018 | ECCV | Deep Bilevel Learning. | Simon Jenni, Paolo Favaro |