| 2025 | WaveFontStyler: Font Style Transfer Based on Sound. | Kota Izumi, Keiji Yanai |
| 2025 | FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation. | Yuki Imajuku, Yoko Yamakata, Kiyoharu Aizawa |
| 2025 | Fingering Prediction for Classical Guitar: Dataset Creation and Model Development. | Nami Iino, Akinaru Iino |
| 2025 | Rotation Methods for 360-Degree Videos in Virtual Reality - A Comparative Study. | Wolfgang Hrst, Leo Zeches |
| 2025 | Innovative Lifelog Visualization and Exploration in Virtual Reality - A Comparative Study. | Wolfgang Hrst, Yannick Visser |
| 2025 | Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models. | Jia-Hong Huang, Hongyi Zhu, Yixian Shen, Stevan Rudinac, Evangelos Kanoulas |
| 2025 | Select and Order: Optimizing Few-Shot Image Classification with In-Context Learning. | Hujiang Huang, Yu Xie, Jun Gao, Chuanliu Fan, Ziqiang Cao |
| 2025 | Flat Local Minima for Continual Learning on Semantic Segmentation. | Zhongzhan Huang, Mingfu Liang, Senwei Liang, Shanshan Zhong |
| 2025 | Movie Retrieval Systems Using Genre-Guided Multimodal Learning Techniques. | Wei-Lun Huang, Shintami Chusnul Hidayati, Tse-Yu Pan |
| 2025 | BLCC: A Benchmark for Multi-LiDAR and Multi-camera Calibration. | Minghui Hou, Gang Wang, Zhiyang Wang, Tongzhou Zhang, Baorui Ma |
| 2025 | SnapSeek 2.0 at Video Browser Showdown 2025. | Minh-Quan Ho-Le, Duy-Khang Ho, Huy-Hoang Do-Huu, Nhut-Thanh Le-Hinh, Hoa-Vien Vo-Hoang, Van-Tu Ninh, Cathal Gurrin, Minh-Triet Tran |
| 2025 | Dynamic Exploration Graph: A Novel Approach for Efficient Nearest Neighbor Search in Evolving Multimedia Datasets. | Nico Hezel, Kai Uwe Barthel, Bruno Schilling, Konstantin Schall, Klaus Jung |
| 2025 | MineTinyNet-YOLO: An Efficient Small Object Detection Method for Complex Underground Coal Mine Scenarios. | Yaling Hao, Wei Wu |
| 2025 | GFA-UDIS: Global-to-Flow Alignment for Unsupervised Deep Image Stitching. | Sijia Han, Zhibin Zhang |
| 2025 | DocMamba: Robust Document Image Dewarping via Selective State Space Sequence Modeling. | Miaolin Han, Huibin Li |
| 2025 | Real-Time Visualizer for Turntablist Performance. | Masatoshi Hamanaka |
| 2025 | MM-CARP: Multimodal Model with Cross-Modal Retrieval-Augmented and Visual Region Perception. | Junhao Guo, Chenhan Fu, Guoming Wang, Rongxing Lu, Dong Chen, Siliang Tang |
| 2025 | Open-Vocabulary Scene Graph Generation via Synonym-Based Predicate Descriptor. | Yuta Goto, Satoshi Yamazaki, Takashi Shibata, Jianquan Liu |
| 2025 | NII-UIT at VBS2025: Multimodal Video Retrieval with LLM Integration and Dynamic Temporal Search. | Bao Tran Gia, Tuong Bui Cong Khanh, Tam Le Thi Thanh, Thuyen Tran Doan, Khiem Le, Tien Do, Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh |
| 2025 | Counting Unique Objects in Geo-Tagged Street Images: A Case Study of Homeless Encampments in Los Angeles. | Narges Ghasemi, Seon Ho Kim, Abdullah Alfarrarjeh, Cyrus Shahabi |
| 2025 | Smart Driving Assistance with Real-Time Risk Assessment and Personalized Driving Coaching to Enhance Road Safety. | Wenbin Gan, Minh-Son Dao, Koji Zettsu |
| 2025 | Can Masking Background and Object Reduce Static Bias for Zero-Shot Action Recognition? | Takumi Fukuzawa, Kensho Hara, Hirokatsu Kataoka, Toru Tamaki |
| 2025 | System Demo of Modeling Smart University Campus Virtual Environments. | Jaime B. Fernandez, Muhammad Intizar Ali |
| 2025 | Structural Information-Guided Fine-Grained Texture Image Inpainting. | Zhiyi Fang, Yi Qian, Xiyue Dai |
| 2025 | HierArtEx: Hierarchical Representations and Art Experts Supporting the Retrieval of Museums in the Metaverse. | Alex Falcon, Ali Abdari, Giuseppe Serra |