| 2026 | OSEG: Improving Diffusion sampling through Orthogonal Smoothed Energy Guidance. | Masud An Nur Islam Fahim, Nazmus Saqib, Joon-Min Gil |
| 2026 | R-MMA: Enhancing Vision-Language Models with Recurrent Adapters for Few-Shot and Cross-Domain Generalization. | Md Fahim, Md Farhan Ishmam, Mir Sazzat Hossain, M. Ashraful Amin, Amin Ahsan Ali, AKM Mahbubur Rahman |
| 2026 | CLIP's Visual Embedding Projector is a Few-shot Cornucopia. | Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick Prez, Raoul de Charette |
| 2026 | An Empirical Study of Siamese Vision Transformers for Scribe Re-Identification. | Alessio Fagioli, Carmine Fabbri, Luigi Cinque, Emanuela Colombi, Gian Luca Foresti |
| 2026 | CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding. | Fevziye Irem Eyiokur, Dogucan Yaman, Hazim Kemal Ekenel, Alexander Waibel |
| 2026 | Conjuring Positive Pairs for Efficient Unification of Representation Learning and Image Synthesis. | Imanol G. Estepa, Jess M. Rodrguez-de-Vera, Ignacio Sarasa, Bhalaji Nagarajan, Petia Radeva |
| 2026 | Direct Visual Grounding by Directing Attention of Visual Tokens. | Parsa Esmaeilkhani, Longin Jan Latecki |
| 2026 | KD360-VoxelBEV: LiDAR and 360-degree Camera Cross Modality Knowledge Distillation for Bird's-Eye-View Segmentation. | Wenke E, Yixin Sun, Jiaxu Liu, Hubert P. H. Shum, Amir Atapour-Abarghouei, Toby P. Breckon |
| 2026 | PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding. | Iaki Erregue, Kamal Nasrollahi, Sergio Escalera |
| 2026 | S3-CLIP: Video Super Resolution for Person-ReID. | Tams Endrei, Gyrgy Cserey |
| 2026 | Zero-LEAD: Source-Free Universal Domain Adaptation for Abdominal Multi-Organ Segmentation. | Ahmed El-Sayed, Marwan Torki |
| 2026 | Deep Image Decomposition for Medical Imaging Anonymization and Curation. | Yael Elkin, Gal Ben-Arie, Tammy Riklin-Raviv |
| 2026 | Isolating the Role of Temporal Information in Video Saliency: A Controlled Experimental Analysis. | Peter El-Jiz, Matthias Kmmerer, Matthias Tangemann, Matthias Bethge, Andreas M. Bartels, Michael M. Bannert |
| 2026 | INRetouch: Context Aware Implicit Neural Representation for Photography Retouching. | Omar Elezabi, Marcos V. Conde, Zongwei Wu, Radu Timofte |
| 2026 | MedProbCLIP: Probabilistic Adaptation of Vision-Language Foundation Model for Reliable Radiograph-Report Retrieval. | Ahmad Elallaf, Yu Zhang, Yuktha Priya Masupalli, Jeong Yang, Young Lee, Zechun Cao, Gongbo Liang |
| 2026 | Towards Consistent and Efficient Decision-Based Attacks. | Henning Duwe, Anna L. Mnz, Holger H. Hoos |
| 2026 | Visibility guided Self-Supervised Occlusion-Resilient Human Pose Estimation. | Arindam Dutta, Sarosij Bose, Rohit Kundu, Calvin-Khang Ta, Saketh Bachu, Konstantinos Karydis, Amit K. Roy-Chowdhury |
| 2026 | Causality-Driven Audits of Model Robustness. | Nathan Drenkow, William Paul, Chris Ribaudo, Mathias Unberath |
| 2026 | Vision Language Models Learn to Assess Images with Specialists. | Quyet V. Do, Seunghyun Yoon, Ruiyi Zhang, Thiloshon Nagarajah, Trung Bui, Viet Dac Lai |
| 2026 | UI-Styler: Ultrasound Image Style Transfer with Class-Aware Prompts for Cross-Device Diagnosis Using a Frozen Black-Box Inference Network. | Nhat-Tuong Do-Tran, Ngoc-Hoang-Lam Le, Ching-Chun Huang |
| 2026 | Towards Inclusive Biometrics: Synthetic Generation of Vitiligo Faces and Their Impact on Face Image Quality. | Andr Drsch, A. Laguna Liang, Christian Rathgeb, Christoph Busch |
| 2026 | VAOT: Vessel-Aware Optimal Transport for Retinal Fundus Enhancement. | Xuanzhao Dong, Wenhui Zhu, Yujian Xiong, Xiwen Chen, Hao Wang, Xin Li, Jiajun Cheng, Zhipeng Wang, Shao Tang, Oana M. Dumitrascu, Yalin Wang |
| 2026 | FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder. | Zeyu Dong, Yimin Zhu, Yu Wu, Yu Sun |
| 2026 | Bridging Restoration and Diagnosis: A Comprehensive Benchmark for Retinal Fundus Enhancement. | Xuanzhao Dong, Wenhui Zhu, Xiwen Chen, Hao Wang, Xin Li, Yujian Xiong, Jiajun Cheng, Zhipeng Wang, Shao Tang, Oana M. Dumitrascu, Yalin Wang |
| 2026 | ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models. | Sibo Dong, Ismail Shaheen, Maggie Shen, Rupayan Mallick, Sarah Adel Bargal |