| 2026 | WACV | VOCAL: Visual Odometry via ContrAstive Learning. | Chi-Yao Huang, Zeel Bhatt, Yezhou Yang |
| 2026 | WACV | Improving Shape Bias in Learnable Geometric Moment Representations. | Sangmin Jung, Anirudh Rayas, Reza Rahimi Azghan, Hassan Ghasemzadeh, Yezhou Yang, Pavan Turaga |
| 2026 | WACV | Event-based Graph Representation with Spatial and Motion Vectors for Asynchronous Object Detection. | Aayush Atul Verma, Arpitsinh Vaghela, Bharatesh Chakravarthi, Kaustav Chanda, Yezhou Yang |
| 2026 | WACV | eSkiTB: A Synthetic Event-Based Dataset for Tracking Skiers. | Krishna Vinod, Joseph Raj Vishal, Kaustav Chanda, Prithvi Jai Ramesh, Yezhou Yang, Bharatesh Chakravarthi |
| 2025 | CVPR | Event Quality Score (EQS): Assessing the Realism of Simulated Event Camera Streams via Distance in Latent Space. | Kaustav Chanda, Aayush Atul Verma, Arpitsinh Vaghela, Yezhou Yang, Bharatesh Chakravarthi |
| 2025 | CVPR | TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark. | Forouzan Fallah, Maitreya Patel, Agneet Chatterjee, Vlad I. Morariu, Chitta Baral, Yezhou Yang |
| 2025 | EMNLP | AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models. | Vatsal Malaviya, Agneet Chatterjee, Maitreya Patel, Yezhou Yang, Chitta Baral |
| 2025 | ICCV | FlowChef: Steering of Rectified Flow Models for Controlled Generations. | Maitreya Patel, Song Wen, Dimitris N. Metaxas, Yezhou Yang |
| 2025 | ICCV | RefEdit: A Benchmark and Method for Improving Instruction-Based Image Editing Model on Referring Expressions. | Bimsara Pathiraja, Maitreya Patel, Shivam Singh, Yezhou Yang, Chitta Baral |
| 2025 | ICLR | VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning. | Nilay Yilmaz, Maitreya Patel, Yiran Lawrence Luo, Tejas Gokhale, Chitta Baral, Suren Jayasuriya, Yezhou Yang |
| 2025 | IJCAI | DeepShade: Enable Shade Simulation by Text-conditioned Image Generation. | Longchao Da, Xiangrui Liu, Mithun Shivakoti, Thirulogasankar Pranav Kutralingam, Yezhou Yang, Hua Wei |
| 2025 | WACV | Deep Geometric Moments Promote Shape Consistency in Text-to-3D Generation. | Utkarsh Nath, Rajeev Goel, Eun Som Jeon, Changhoon Kim, Kyle Min, Yezhou Yang, Yingzhen Yang, Pavan K. Turaga |
| 2024 | AAAI | ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models. | Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang |
| 2024 | CVPR | On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation. | Agneet Chatterjee, Tejas Gokhale, Chitta Baral, Yezhou Yang |
| 2024 | CVPR | 'Eyes of a Hawk and Ears of a Fox': Part Prototype Network for Generalized Zero-Shot Learning. | Joshua Feinglass, Jayaraman J. Thiagarajan, Rushil Anirudh, T. S. Jayram, Yezhou Yang |
| 2024 | CVPR | WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models. | Changhoon Kim, Kyle Min, Maitreya Patel, Sheng Cheng, Yezhou Yang |
| 2024 | CVPR | Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images. | Yiran Luo, Joshua Feinglass, Tejas Gokhale, Kuan-Cheng Lee, Chitta Baral, Yezhou Yang |
| 2024 | CVPR | ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations. | Maitreya Patel, Changhoon Kim, Sheng Cheng, Chitta Baral, Yezhou Yang |
| 2024 | CVPR | eTraM: Event-Based Traffic Monitoring Dataset. | Aayush Atul Verma, Bharatesh Chakravarthi, Arpitsinh Vaghela, Hua Wei, Yezhou Yang |
| 2024 | CVPR | Evaluating Multimodal Large Language Models across Distribution Shifts and Augmentations. | Aayush Atul Verma, Amir Saeidi, Shamanthak Hegde, Ajay Therala, Fenil Denish Bardoliya, Nagaraju Machavarapu, Shri Ajay Kumar Ravindhiran, Srija Malyala, Agneet Chatterjee, Yezhou Yang, Chitta Baral |
| 2024 | ECCV | Recent Event Camera Innovations: A Survey. | Bharatesh Chakravarthi, Aayush Atul Verma, Kostas Daniilidis, Cornelia Fermller, Yezhou Yang |
| 2024 | ECCV | REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models. | Agneet Chatterjee, Yiran Luo, Tejas Gokhale, Yezhou Yang, Chitta Baral |
| 2024 | ECCV | Getting it Right: Improving Spatial Consistency in Text-to-Image Models. | Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo, Sayak Paul, Dhruba Ghosh, Tejas Gokhale, Ludwig Schmidt, Hannaneh Hajishirzi, Vasudev Lal, Chitta Baral, Yezhou Yang |
| 2024 | ECCV | R.A.C.E. : Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model. | Changhoon Kim, Kyle Min, Yezhou Yang |
| 2024 | EMNLP | Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model. | Sheng Cheng, Maitreya Patel, Yezhou Yang |
| 2024 | EMNLP | TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning. | Joshua Feinglass, Yezhou Yang |
| 2024 | NAACL | Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts. | Michael Saxon, Yiran Luo, Sharon Levy, Chitta Baral, Yezhou Yang, William Yang Wang |
| 2024 | WACV | Towards Addressing the Misalignment of Object Proposal Evaluation for Vision-Language Tasks via Semantic Grounding. | Joshua Feinglass, Yezhou Yang |
| 2023 | ACL | End-to-end Knowledge Retrieval with Multi-modal Queries. | Man Luo, Zhiyuan Fang, Tejas Gokhale, Yezhou Yang, Chitta Baral |
| 2023 | ICCV | Adversarial Bayesian Augmentation for Single-Source Domain Generalization. | Sheng Cheng, Tejas Gokhale, Yezhou Yang |
| 2023 | ICML | Attributing Image Generative Models using Latent Fingerprints. | Guangyu Nie, Changhoon Kim, Yezhou Yang, Yi Ren |
| 2023 | ICRA | CAROM Air - Vehicle Localization and Traffic Scene Reconstruction from Aerial Videos. | Duo Lu, Eric Eaton, Matt Weg, Wei Wang, Steven Como, Jeffrey Wishart, Hongbin Yu, Yezhou Yang |
| 2023 | WACV | Improving Diversity with Adversarially Learned Transformations for Domain Generalization. | Tejas Gokhale, Rushil Anirudh, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Chitta Baral, Yezhou Yang |
| 2022 | ACL | Semantically Distributed Robust Optimization for Vision-and-Language Inference. | Tejas Gokhale, Abhishek Chaudhary, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
| 2022 | ACL | To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo. | Yiran Luo, Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta Baral |
| 2022 | CVPR | Good, Better, Best: Textual Distractors Generation for Multiple-Choice Visual Question Answering via Reinforcement Learning. | Jiaying Lu, Xin Ye, Yi Ren, Yezhou Yang |
| 2022 | CVPR | Tragedy Plus Time: Capturing Unintended Human Activities from Weakly-labeled Videos. | Arnav Chakravarthy, Zhiyuan Fang, Yezhou Yang |
| 2022 | CVPR | SSR-GNNs: Stroke-based Sketch Representation with Graph Neural Networks. | Sheng Cheng, Yi Ren, Yezhou Yang |
| 2022 | CVPR | Injecting Semantic Concepts into End-to-End Image Captioning. | Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang, Zhe Gan, Lijuan Wang, Yezhou Yang, Zicheng Liu |
| 2022 | EMNLP | CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering. | Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang |
| 2022 | EMNLP | Learning Action-Effect Dynamics for Hypothetical Vision-Language Reasoning Task. | Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang, Chitta Baral |
| 2022 | ICASSP | Attributable Watermarking of Speech Generative Models. | Yongbaek Cho, Changhoon Kim, Yezhou Yang, Yi Ren |
| 2022 | ICPR | CAVAN: Commonsense Knowledge Anchored Video Captioning. | Huiliang Shao, Zhiyuan Fang, Yezhou Yang |
| 2022 | ICRA | Targeted Attack on Deep RL-based Autonomous Driving with Learned Visual Patterns. | Prasanth Buddareddygari, Travis Zhang, Yezhou Yang, Yi Ren |
| 2021 | AAAI | Attribute-Guided Adversarial Training for Robustness to Natural Perturbations. | Tejas Gokhale, Rushil Anirudh, Bhavya Kailkhura, Jayaraman J. Thiagarajan, Chitta Baral, Yezhou Yang |
| 2021 | ACL | WeaQA: Weak Supervision via Captions for Visual Question Answering. | Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta Baral |
| 2021 | ACL | SMURF: SeMantic and linguistic UndeRstanding Fusion for Caption Evaluation via Typicality Analysis. | Joshua Feinglass, Yezhou Yang |
| 2021 | CIKM | First Workshop on Knowledge Injection in Neural Networks (KINN). | Vasudev Lal, Somak Aditya, Yezhou Yang, Pasquale Minervini, Sandya Mannarswamy |
| 2021 | CVPR | Hierarchical and Partially Observable Goal-Driven Policy Learning With Goals Relational Graph. | Xin Ye, Yezhou Yang |
| 2021 | ICCV | Weakly Supervised Relative Spatial Reasoning for Visual Question Answering. | Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta Baral |
| 2021 | ICCV | Compressing Visual-linguistic Model via Knowledge Distillation. | Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lijuan Wang, Yezhou Yang, Zicheng Liu |
| 2021 | ICLR | SEED: Self-supervised Distillation For Visual Representation. | Zhiyuan Fang, Jianfeng Wang, Lijuan Wang, Lei Zhang, Yezhou Yang, Zicheng Liu |
| 2021 | ICLR | Decentralized Attribution of Generative Models. | Changhoon Kim, Yi Ren, Yezhou Yang |
| 2021 | ICRA | CAROM - Vehicle Localization and Traffic Scene Reconstruction from Monocular Cameras on Road Infrastructures. | Duo Lu, Varun C. Jammula, Steven Como, Jeffrey Wishart, Yan Chen, Yezhou Yang |
| 2021 | NAACL | CLEVR_HYP: A Challenge Dataset and Baselines for Visual Question Answering with Hypothetical Actions over Images. | Shailaja Keyur Sampat, Akshay Kumar, Yezhou Yang, Chitta Baral |
| 2020 | CVPR | Enabling Incremental Knowledge Transfer for Object Detection at the Edge. | Mohammad Farhadi, Mehdi Ghasemi, Sarma B. K. Vrudhula, Yezhou Yang |
| 2020 | ECCV | VQA-LOL: Visual Question Answering Under the Lens of Logic. | Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
| 2020 | ECCV | ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language. | Zhe Wang, Zhiyuan Fang, Jun Wang, Yezhou Yang |
| 2020 | EMNLP | Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning. | Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
| 2020 | EMNLP | MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question Answering. | Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
| 2020 | EMNLP | Visuo-Lingustic Question Answering (VLQA) Challenge. | Shailaja Keyur Sampat, Yezhou Yang, Chitta Baral |
| 2020 | IROS | Learning hierarchical behavior and motion planning for autonomous driving. | Jingke Wang, Yue Wang, Dongkun Zhang, Yezhou Yang, Rong Xiong |
| 2020 | WACV | TKD: Temporal Knowledge Distillation for Active Perception. | Mohammad Farhadi, Yezhou Yang |
| 2019 | CVPR | Modularized Textual Grounding for Counterfactual Resilience. | Zhiyuan Fang, Shu Kong, Charless C. Fowlkes, Yezhou Yang |
| 2019 | CVPR | Cooking With Blocks : A Recipe for Visual Reasoning on Image-Pairs. | Tejas Gokhale, Shailaja Sampat, Zhiyuan Fang, Yezhou Yang, Chitta Baral |
| 2019 | ICIP | Image Decomposition and Classification Through a Generative Model. | Houpu Yao, Malcolm Regan, Yezhou Yang, Yi Ren |
| 2019 | IJCAI | Integrating Knowledge and Reasoning in Image Understanding. | Somak Aditya, Yezhou Yang, Chitta Baral |
| 2019 | ICRA | How Shall I Drive? Interaction Modeling and Motion Planning towards Empathetic and Socially-Graceful Driving. | Yi Ren, Steven Elliott, Yiwei Wang, Yezhou Yang, Wenlong Zhang |
| 2019 | WACV | Spatial Knowledge Distillation to Aid Visual Reasoning. | Somak Aditya, Rudra Saha, Yezhou Yang, Chitta Baral |
| 2018 | AAAI | Explicit Reasoning over End-to-End Neural Architectures for Visual Question Answering. | Somak Aditya, Yezhou Yang, Chitta Baral |
| 2018 | CVPR | Transductive Unbiased Embedding for Zero-Shot Learning. | Jie Song, Chengchao Shen, Yezhou Yang, Yang Liu, Mingli Song |
| 2018 | ECCV | Stroke Controllable Fast Style Transfer with Adaptive Receptive Fields. | Yongcheng Jing, Yang Liu, Yezhou Yang, Zunlei Feng, Yizhou Yu, Dacheng Tao, Mingli Song |
| 2018 | ICIP | DeepSSH: Deep Semantic Structured Hashing for Explainable Person Re-Identification. | Ya Zhao, Sihui Luo, Yezhou Yang, Mingli Song |
| 2018 | ICONIP | DeepSIC: Deep Semantic Image Compression. | Sihui Luo, Yezhou Yang, Yanling Yin, Chengchao Shen, Ya Zhao, Mingli Song |
| 2018 | IROS | Active Object Perceiver: Recognition-Guided Policy Learning for Object Searching on Mobile Robots. | Xin Ye, Zhe Lin, Haoxiang Li, Shibin Zheng, Yezhou Yang |
| 2018 | ICRA | Extrinsic Dexterity Through Active Slip Control Using Deep Predictive Models. | Simon Stepputtis, Yezhou Yang, Heni Ben Amor |
| 2018 | MICCAI | Weakly-Supervised Learning-Based Feature Localization for Confocal Laser Endomicroscopy Glioma Images. | Mohammadhassan Izadyyazdanabadi, Evgenii Belykh, Claudio Cavallo, Xiaochun Zhao, Sirin Gandhi, Leandro Borba Moreira, Jennifer Eschbacher, Peter Nakaji, Mark C. Preul, Yezhou Yang |
| 2018 | UAI | Combining Knowledge and Reasoning through Probabilistic Soft Logic for Image Puzzle Solving. | Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos |
| 2017 | CVPR | Hand Movement Prediction Based Collision-Free Human-Robot Interaction. | Yiwei Wang, Xin Ye, Yezhou Yang, Wenlong Zhang |
| 2017 | ICRA | Fast task-specific target detection via graph based constraints representation and checking. | Wentao Luan, Yezhou Yang, Cornelia Fermller, John S. Baras |
| 2017 | ICRA | What can i do around here? Deep functional scene understanding for cognitive robots. | Chengxi Ye, Yezhou Yang, Ren Mao, Cornelia Fermller, Yiannis Aloimonos |
| 2016 | ECCV | Reliable Attribute-Based Object Recognition Using High Predictive Value Classifiers. | Wentao Luan, Yezhou Yang, Cornelia Fermller, John S. Baras |
| 2015 | AAAI | Robot Learning Manipulation Action Plans by "Watching" Unconstrained Videos from the World Wide Web. | Yezhou Yang, Yi Li, Cornelia Fermller, Yiannis Aloimonos |
| 2015 | ACL | Learning the Semantics of Manipulation Action. | Yezhou Yang, Yiannis Aloimonos, Cornelia Fermller, Eren Erdal Aksoy |
| 2015 | CVPR | Grasp type revisited: A modern perspective on a classical feature for vision. | Yezhou Yang, Cornelia Fermller, Yi Li, Yiannis Aloimonos |
| 2015 | ICRA | Learning the spatial semantics of manipulation actions through preposition grounding. | Konstantinos Zampogiannis, Yezhou Yang, Cornelia Fermller, Yiannis Aloimonos |
| 2013 | BMVC | Action Attribute Detection from Sports Videos with Contextual Constraints. | Xiaodong Yu, Ching Lik Teo, Yezhou Yang, Cornelia Fermller, Yiannis Aloimonos |
| 2013 | CVPR | Detection of Manipulation Action Consequences (MAC). | Yezhou Yang, Cornelia Fermller, Yiannis Aloimonos |
| 2013 | IROS | Minimalist plans for interpreting manipulation actions. | Anupam Guha, Yezhou Yang, Cornelia Fermller, Yiannis Aloimonos |
| 2013 | ICRA | Robots with language: Multi-label visual recognition using NLP. | Yezhou Yang, Ching Lik Teo, Cornelia Fermller, Yiannis Aloimonos |
| 2012 | IROS | Using a minimal action grammar for activity understanding in the real world. | Douglas Summers-Stay, Ching Lik Teo, Yezhou Yang, Cornelia Fermller, Yiannis Aloimonos |
| 2012 | ICRA | Towards a Watson that sees: Language-guided action recognition for robots. | Ching Lik Teo, Yezhou Yang, Hal Daum III, Cornelia Fermller, Yiannis Aloimonos |
| 2011 | AAAI | A Corpus-Guided Framework for Robotic Visual Perception. | Ching Lik Teo, Yezhou Yang, Hal Daum III, Cornelia Fermller, Yiannis Aloimonos |
| 2011 | EMNLP | Corpus-Guided Sentence Generation of Natural Images. | Yezhou Yang, Ching Lik Teo, Hal Daum III, Yiannis Aloimonos |
| 2011 | ICCV | Active scene recognition with vision and language. | Xiaodong Yu, Cornelia Fermller, Ching Lik Teo, Yezhou Yang, Yiannis Aloimonos |
| 2010 | ECCV | What Is the Chance of Happening: A New Way to Predict Where People Look. | Yezhou Yang, Mingli Song, Na Li, Jiajun Bu, Chun Chen |