| 2026 | ACL | Agentic Very Long Video Understanding. | Aniket Rege, Arka Sadhu, Yuliang Li, Kejie Li, Ramya Korlakai Vinayak, Yuning Chai, Yong Jae Lee, Hyo Jin Kim |
| 2026 | SIGGRAPH | MAOAM: Unified Object and Material Selection with Vision-Language Models. | Jaden Park, Valentin Deschaintre, Jason Kuen, Kangning Liu, Iliyan Georgiev, Krishna Kumar Singh, Yong Jae Lee, Michael Fischer |
| 2026 | WACV | LASER: Lip Landmark Assisted Speaker Detection for Robustness. | Le Thien Phuc Nguyen, Zhuoran Yu, Yong Jae Lee |
| 2025 | CVPR | Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs. | Zeyi Huang, Yuyang Ji, Xiaofang Wang, Nikhil Mehta, Tong Xiao, Donghyun Lee, Sigmund Vanvalkenburgh, Shengxin Zha, Bolin Lai, Licheng Yu, Ning Zhang, Yong Jae Lee, Miao Liu |
| 2025 | CVPR | Yo'Chameleon: Personalized Vision and Language Generation. | Thao Nguyen, Krishna Kumar Singh, Jing Shi, Trung Bui, Yong Jae Lee, Yuheng Li |
| 2025 | ICCV | Customizing Domain Adapters for Domain Generalization. | Yuyang Ji, Zeyi Huang, Haohan Wang, Yong Jae Lee |
| 2025 | ICCV | X-Fusion: Introducing New Modality to Frozen Large Language Models. | Sicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer, Yijun Li, Yuchen Liu, Abhishek Tandon, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou, Yuheng Li |
| 2025 | ICCV | CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems. | Aniket Rege, Zinnia Nie, Mahesh Ramesh, Unmesh Raskar, Zhuoran Yu, Aditya Kusupati, Yong Jae Lee, Ramya Korlakai Vinayak |
| 2025 | ICCV | LLaVA-Prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models. | Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, Yan Yan |
| 2025 | ICLR | LLaRA: Supercharging Robot Learning Data for Vision-Language Policy. | Xiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan D. Burgert, Mu Cai, Yong Jae Lee, Michael S. Ryoo |
| 2025 | ICLR | Matryoshka Multimodal Models. | Mu Cai, Jianwei Yang, Jianfeng Gao, Yong Jae Lee |
| 2025 | ICLR | Aligned Datasets Improve Detection of Latent Diffusion-Generated Images. | Anirudh Sundara Rajan, Utkarsh Ojha, Jedidiah Schloesser, Yong Jae Lee |
| 2025 | ICML | Stay-Positive: A Case for Ignoring Real Image Features in Fake Image Detection. | Anirudh Sundara Rajan, Yong Jae Lee |
| 2025 | ICRA | Cohere3D: Exploiting Temporal Coherence for Unsupervised Representation Learning of Vision-Based Autonomous Driving. | Yichen Xie, Hongge Chen, Gregory P. Meyer, Yong Jae Lee, Eric M. Wolff, Masayoshi Tomizuka, Wei Zhan, Yuning Chai, Xin Huang |
| 2025 | WACV | An Investigation on LLMs' Visual Understanding Ability Using SVG for Image-Text Bridging. | Mu Cai, Zeyi Huang, Yuheng Li, Utkarsh Ojha, Haohan Wang, Yong Jae Lee |
| 2024 | ACL | CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples. | Jianrui Zhang, Mu Cai, Tengyang Xie, Yong Jae Lee |
| 2024 | CVPR | ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts. | Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, Yong Jae Lee |
| 2024 | CVPR | Improved Baselines with Visual Instruction Tuning. | Haotian Liu, Chunyuan Li, Yuheng Li, Yong Jae Lee |
| 2024 | CVPR | Edit One for All: Interactive Batch Image Editing. | Thao Nguyen, Utkarsh Ojha, Yuheng Li, Haotian Liu, Yong Jae Lee |
| 2024 | ECCV | Removing Distributional Discrepancies in Captions Improves Image-Text Alignment. | Yuheng Li, Haotian Liu, Mu Cai, Yijun Li, Eli Shechtman, Zhe Lin, Yong Jae Lee, Krishna Kumar Singh |
| 2024 | EMNLP | MATE: Meet At The Embedding - Connecting Images with Long Texts. | Young Kyun Jang, Junmo Kang, Yong Jae Lee, Donghyun Kim |
| 2024 | EMNLP | VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation. | Bocheng Zou, Mu Cai, Jianrui Zhang, Yong Jae Lee |
| 2024 | IROS | Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds. | Mu Cai, Chenxu Luo, Yong Jae Lee, Xiaodong Yang |
| 2024 | WACV | Computer Vision on the Edge: Individual Cattle Identification in Real-time with ReadMyCow System. | Moniek Smink, Haotian Liu, Drte Dpfer, Yong Jae Lee |
| 2023 | CVPR | GLIGEN: Open-Set Grounded Text-to-Image Generation. | Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, Yong Jae Lee |
| 2023 | CVPR | Learning Customized Visual Models with Retrieval-Augmented Knowledge. | Haotian Liu, Kilho Son, Jianwei Yang, Ce Liu, Jianfeng Gao, Yong Jae Lee, Chunyuan Li |
| 2023 | CVPR | Towards Universal Fake Image Detectors that Generalize Across Generative Models. | Utkarsh Ojha, Yuheng Li, Yong Jae Lee |
| 2023 | CVPR | Generalized Decoding for Pixel, Image, and Language. | Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, Nanyun Peng, Lijuan Wang, Yong Jae Lee, Jianfeng Gao |
| 2023 | ICCV | A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language Guidance. | Zeyi Huang, Andy Zhou, Zijian Lin, Mu Cai, Haohan Wang, Yong Jae Lee |
| 2023 | ICLR | InPL: Pseudo-labeling the Inliers First for Imbalanced Semi-supervised Learning. | Zhuoran Yu, Yin Li, Yong Jae Lee |
| 2022 | CVPR | The Two Dimensions of Worst-case Training and Their Integrated Effect for Out-of-domain Generalization. | Zeyi Huang, Haohan Wang, Dong Huang, Yong Jae Lee, Eric P. Xing |
| 2022 | CVPR | GIRAFFE HD: A High-Resolution 3D-aware Generative Model. | Yang Xue, Yuheng Li, Krishna Kumar Singh, Yong Jae Lee |
| 2022 | ECCV | Contrastive Learning for Diverse Disentangled Foreground Generation. | Yuheng Li, Yijun Li, Jingwan Lu, Eli Shechtman, Yong Jae Lee, Krishna Kumar Singh |
| 2022 | ECCV | Masked Discrimination for Self-supervised Learning on Point Clouds. | Haotian Liu, Mu Cai, Yong Jae Lee |
| 2022 | WACV | Equine Pain Behavior Classification via Self-Supervised Disentangled Pose Representation. | Maheen Rashid, Sofia Broom, Katrina Ask, Elin Hernlund, Pia Haubro Andersen, Hedvig Kjellstrm, Yong Jae Lee |
| 2022 | UAI | Toward learning human-aligned cross-domain robust models by countering misaligned features. | Haohan Wang, Zeyi Huang, Hanlin Zhang, Yong Jae Lee, Eric P. Xing |
| 2021 | BMVC | PartGAN: Unsupervised Part Decomposition for Image Generation and Segmentation. | Yuheng Li, Krishna Kumar Singh, Yang Xue, Yong Jae Lee |
| 2021 | CVPR | Few-Shot Image Generation via Cross-Domain Correspondence. | Utkarsh Ojha, Yijun Li, Jingwan Lu, Alexei A. Efros, Yong Jae Lee, Eli Shechtman, Richard Zhang |
| 2021 | CVPR | Progressive Temporal Feature Alignment Network for Video Inpainting. | Xueyan Zou, Linjie Yang, Ding Liu, Yong Jae Lee |
| 2021 | ICCV | Collaging Class-specific GANs for Semantic Image Synthesis. | Yuheng Li, Yijun Li, Jingwan Lu, Eli Shechtman, Yong Jae Lee, Krishna Kumar Singh |
| 2021 | ICLR | Generating Furry Cars: Disentangling Object Shape and Appearance across Multiple Domains. | Utkarsh Ojha, Krishna Kumar Singh, Yong Jae Lee |
| 2021 | ICRA | YolactEdge: Real-time Instance Segmentation on the Edge. | Haotian Liu, Rafael A. Rivera Soto, Fanyi Xiao, Yong Jae Lee |
| 2021 | WACV | SinGAN-GIF: Learning a Generative Video Model from a Single GIF. | Rajat Arora, Yong Jae Lee |
| 2020 | BMVC | Delving Deeper into Anti-aliasing in ConvNets. | Xueyan Zou, Fanyi Xiao, Zhiding Yu, Yong Jae Lee |
| 2020 | CVPR | MixNMatch: Multifactor Disentanglement and Encoding for Conditional Image Generation. | Yuheng Li, Krishna Kumar Singh, Utkarsh Ojha, Yong Jae Lee |
| 2020 | CVPR | Instance-Aware, Context-Focused, and Memory-Efficient Weakly Supervised Object Detection. | Zhongzheng Ren, Zhiding Yu, Xiaodong Yang, Ming-Yu Liu, Yong Jae Lee, Alexander G. Schwing, Jan Kautz |
| 2020 | CVPR | Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias. | Krishna Kumar Singh, Dhruv Mahajan, Kristen Grauman, Yong Jae Lee, Matt Feiszli, Deepti Ghadiyaram |
| 2020 | ECCV | Password-Conditioned Anonymization and Deanonymization with Face Identity Transformers. | Xiuye Gu, Weixin Luo, Michael S. Ryoo, Yong Jae Lee |
| 2020 | WACV | Action Graphs: Weakly-supervised Action Localization with Graph Convolution Networks. | Maheen Rashid, Hedvig Kjellstrm, Yong Jae Lee |
| 2019 | CVPR | HPLFlowNet: Hierarchical Permutohedral Lattice FlowNet for Scene Flow Estimation on Large-Scale Point Clouds. | Xiuye Gu, Yijie Wang, Chongruo Wu, Yong Jae Lee, Panqu Wang |
| 2019 | CVPR | You Reap What You Sow: Using Videos to Generate High Precision Object Proposals for Weakly-Supervised Object Detection. | Krishna Kumar Singh, Yong Jae Lee |
| 2019 | CVPR | FineGAN: Unsupervised Hierarchical Disentanglement for Fine-Grained Object Generation and Discovery. | Krishna Kumar Singh, Utkarsh Ojha, Yong Jae Lee |
| 2019 | ICCV | YOLACT: Real-Time Instance Segmentation. | Daniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae Lee |
| 2019 | ICCV | Identity From Here, Pose From There: Self-Supervised Disentanglement and Generation of Objects Using Unlabeled Videos. | Fanyi Xiao, Haotian Liu, Yong Jae Lee |
| 2018 | CVPR | Cross-Domain Self-Supervised Multi-Task Feature Learning Using Synthetic Imagery. | Zhongzheng Ren, Yong Jae Lee |
| 2018 | ECCV | Learning to Anonymize Faces for Privacy Preserving Action Detection. | Zhongzheng Ren, Yong Jae Lee, Michael S. Ryoo |
| 2018 | ECCV | DOCK: Detecting Objects by Transferring Common-Sense Knowledge. | Krishna Kumar Singh, Santosh Kumar Divvala, Ali Farhadi, Yong Jae Lee |
| 2018 | ECCV | Video Object Detection with an Aligned Spatial-Temporal Memory. | Fanyi Xiao, Yong Jae Lee |
| 2018 | EMNLP | A Visual Attention Grounding Neural Model for Multimodal Machine Translation. | Mingyang Zhou, Runxiang Cheng, Yong Jae Lee, Zhou Yu |
| 2018 | WSDM | Who Will Share My Image?: Predicting the Content Diffusion Path in Online Social Networks. | Wenjian Hu, Krishna Kumar Singh, Fanyi Xiao, Jinyoung Han, Chen-Nee Chuah, Yong Jae Lee |
| 2017 | CVPR | Identifying First-Person Camera Wearers in Third-Person Videos. | Chenyou Fan, Jangwon Lee, Mingze Xu, Krishna Kumar Singh, Yong Jae Lee, David J. Crandall, Michael S. Ryoo |
| 2017 | CVPR | Interspecies Knowledge Transfer for Facial Keypoint Detection. | Maheen Rashid, Xiuye Gu, Yong Jae Lee |
| 2017 | CVPR | Weakly-Supervised Visual Grounding of Phrases with Linguistic Structures. | Fanyi Xiao, Leonid Sigal, Yong Jae Lee |
| 2017 | ICCV | Hide-and-Seek: Forcing a Network to be Meticulous for Weakly-Supervised Object and Action Localization. | Krishna Kumar Singh, Yong Jae Lee |
| 2017 | WACV | Who Moved My Cheese? Automatic Annotation of Rodent Behaviors with Convolutional Neural Networks. | Zhongzheng Ren, Adriana Noronha Annie, Vogel Ciernia, Yong Jae Lee |
| 2016 | CVPR | Track and Transfer: Watching Videos to Simulate Strong Human Supervision for Weakly-Supervised Object Detection. | Krishna Kumar Singh, Fanyi Xiao, Yong Jae Lee |
| 2016 | CVPR | Track and Segment: An Iterative Unsupervised Approach for Video Object Proposals. | Fanyi Xiao, Yong Jae Lee |
| 2016 | ECCV | End-to-End Localization and Ranking for Relative Attributes. | Krishna Kumar Singh, Yong Jae Lee |
| 2015 | CVPR | FlowWeb: Joint image set alignment by weaving consistent, pixel-wise correspondences. | Tinghui Zhou, Yong Jae Lee, Stella X. Yu, Alexei A. Efros |
| 2015 | ICCV | Discovering the Spatial Extent of Relative Attributes. | Fanyi Xiao, Yong Jae Lee |
| 2014 | CVPR | An Introduction to the 3rd Workshop on Egocentric (First-Person) Vision. | Steve Mann, Kris M. Kitani, Yong Jae Lee, Michael S. Ryoo, Alireza Fathi |
| 2013 | ICCV | Style-Aware Mid-level Representation for Discovering Visual Connections in Space and Time. | Yong Jae Lee, Alexei A. Efros, Martial Hebert |
| 2012 | CVPR | Discovering important people and objects for egocentric video summarization. | Yong Jae Lee, Joydeep Ghosh, Kristen Grauman |
| 2011 | BMVC | Face Discovery with Social Context. | Yong Jae Lee, Kristen Grauman |
| 2011 | CVPR | Learning the easy things first: Self-paced visual category discovery. | Yong Jae Lee, Kristen Grauman |
| 2011 | ICCV | Key-segments for video object segmentation. | Yong Jae Lee, Jaechul Kim, Kristen Grauman |
| 2010 | CVPR | Object-graphs for context-aware category discovery. | Yong Jae Lee, Kristen Grauman |
| 2010 | CVPR | Collect-cut: Segmentation with top-down cues discovered in multi-object images. | Yong Jae Lee, Kristen Grauman |
| 2009 | CVPR | Shape discovery from unlabeled image collections. | Yong Jae Lee, Kristen Grauman |
| 2008 | BMVC | Foreground Focus: Finding Meaningful Features in Unlabeled Images. | Yong Jae Lee, Kristen Grauman |