| 2025 | Continuous Locomotive Crowd Behavior Generation. | Inhwan Bae, Junoh Lee, Hae-Gon Jeon |
| 2025 | MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations. | Kyungho Bae, Jinhyung Kim, Sihaeng Lee, Soonyoung Lee, Gunhee Lee, Jinwoo Choi |
| 2025 | TADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task Learning. | Seungmin Baek, Soyul Lee, Hayeon Jo, Hyesong Choi, Dongbo Min |
| 2025 | Three Cars Approaching within 100m! Enhancing Distant Geometry by Tri-Axis Voxel Scanning for Camera-based Semantic Scene Completion. | Jongseong Bae, Junwoo Ha, Ha Young Kim |
| 2025 | Robustness Evaluation for Video Models with Reinforcement Learning. | Ashwin Ramesh Babu, Sajad Mousavi, Vineet Gundecha, Sahand Ghorbanpour, Avisek Naug, Antonio Guillen, Ricardo Luna, Soumyendu Sarkar |
| 2025 | Coordinated Robustness Evaluation Framework for Vision-Language Models. | Ashwin Ramesh Babu, Sajad Mousavi, Vineet Gundecha, Sahand Ghorbanpour, Avisek Naug, Antonio Guillen, Ricardo Luna, Soumyendu Sarkar |
| 2025 | Self-Supervised Cross-View Correspondence with Predictive Cycle Consistency. | Alan Baade, Changan Chen |
| 2025 | Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control. | Basim Azam, Naveed Akhtar |
| 2025 | HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding. | Shehreen Azad, Vibhav Vineet, Yogesh Singh Rawat |
| 2025 | Understanding Depth and Height Perception in Large Visual-Language Models. | Shehreen Azad, Yash Jain, Rishit Garg, Vibhav Vineet, Yogesh S. Rawat |
| 2025 | Physics-based Human Pose Estimation from a Single Moving RGB Camera. | Ayce Idil Aytekin, Chuqiao Li, Diogo C. Luvizon, Rishabh Dabral, Martin R. Oswald, Marc Habermann, Christian Theobalt |
| 2025 | ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements. | M. Arda Aydin, Efe Mert irpar, Elvin Abdinli, Gozde Unal, Yusuf Hseyin Sahin |
| 2025 | MTevent: A Multi-Task Event Camera Dataset for 6D Pose Estimation and Moving Object Detection. | Shrutarv Awasthi, Anas Gouda, Sven Franke, Jrme Rutinowski, Frank Hoffmann, Moritz Roidl |
| 2025 | Stable Flow: Vital Layers for Training-Free Image Editing. | Omri Avrahami, Or Patashnik, Ohad Fried, Egor Nemchinov, Kfir Aberman, Dani Lischinski, Daniel Cohen-Or |
| 2025 | PETAH: Parameter Efficient Task Adaptation for Hybrid Transformers. | Maximilian Augustin, Syed Shakib Sarwar, Mostafa Elhoushi, Yuecheng Li, Sai Qian Zhang, Barbara De Salvo |
| 2025 | ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation. | Ali Athar, Xueqing Deng, Liang-Chieh Chen |
| 2025 | Real-Time Ultra-Fine-Grained Surgical Instrument Classification. | Md. Atabuzzaman, Gino DiMatteo, Hani Alomari, Chiawei Tang, Connor Hale, Adam E. Goode, David Ryan King, Chris Thomas |
| 2025 | AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities. | Guillaume Astruc, Nicolas Gonthier, Clment Mallet, Loc Landrieu |
| 2025 | Dense Match Summarization for Faster Two-view Estimation. | Jonathan Astermark, Anders Heyden, Viktor Larsson |
| 2025 | DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation. | Maregu Assefa, Muzammal Naseer, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Mohamed L. Seghier, Naoufel Werghi |
| 2025 | FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs. | Mothilal Asokan, Kebin Wu, Fatima Albreiki |
| 2025 | Maximizing aerial detection of organic objects in non-exhaustively searchable survey area. | Amir Ehsan Niaraki Asli, Jansel Herrera-Gerena, Jeremy Roghair, Ali Jannesari |
| 2025 | Balancing Privacy and Action Performance: A Penalty-Driven Approach to Image Anonymization. | Nazia Aslam, Kamal Nasrollahi |
| 2025 | MET3R: Measuring Multi-View Consistency in Generated Images. | Mohammad Asim, Christopher Wewer, Thomas Wimmer, Bernt Schiele, Jan Eric Lenssen |
| 2025 | FIction: 4D Future Interaction Prediction from Video. | Kumar Ashutosh, Georgios Pavlakos, Kristen Grauman |