| 2024 | Long-Term Temporal Context Gathering for Neural Video Compression. | Linfeng Qi, Zhaoyang Jia, Jiahao Li, Bin Li, Houqiang Li, Yan Lu |
| 2024 | Layered Rendering Diffusion Model for Controllable Zero-Shot Image Synthesis. | Zipeng Qi, Guoxi Huang, Chenyang Liu, Fei Ye |
| 2024 | SignGen: End-to-End Sign Language Video Generation with Latent Diffusion. | Fan Qi, Yu Duan, Huaiwen Zhang, Changsheng Xu |
| 2024 | ShapeLLM: Universal 3D Object Understanding for Embodied Interaction. | Zekun Qi, Runpei Dong, Shaochen Zhang, Haoran Geng, Chunrui Han, Zheng Ge, Li Yi, Kaisheng Ma |
| 2024 | Modelling Competitive Behaviors in Autonomous Driving Under Generative World Model. | Guanren Qiao, Guorui Quan, Rongxiao Qu, Guiliang Liu |
| 2024 | LLM as Copilot for Coarse-Grained Vision-and-Language Navigation. | Yanyuan Qiao, Qianyi Liu, Jiajun Liu, Jing Liu, Qi Wu |
| 2024 | SeA: Semantic Adversarial Augmentation for Last Layer Features from Unsupervised Representation Learning. | Qi Qian, Yuanhong Xu, Juhua Hu |
| 2024 | Text Motion Translator: A Bi-directional Model for Enhanced 3D Human Motion Generation from Open-Vocabulary Descriptions. | Yijun Qian, Jack Urbanek, Alex Hauptmann, Jungdam Won |
| 2024 | Multi-branch Collaborative Learning Network for 3D Visual Grounding. | Zhipeng Qian, Yiwei Ma, Zhekai Lin, Jiayi Ji, Xiawu Zheng, Xiaoshuai Sun, Rongrong Ji |
| 2024 | Online Zero-Shot Classification with CLIP. | Qi Qian, Juhua Hu |
| 2024 | Fairness-Aware Vision Transformer via Debiased Self-Attention. | Yao Qiang, Chengyin Li, Prashant Khanduri, Dongxiao Zhu |
| 2024 | Rethinking Image-to-Video Adaptation: An Object-Centric Perspective. | Rui Qian, Shuangrui Ding, Dahua Lin |
| 2024 | Efficient Diffusion Transformer with Step-Wise Dynamic Attention Mediators. | Yifan Pu, Zhuofan Xia, Jiayi Guo, Dongchen Han, Qixiu Li, Duo Li, Yuhui Yuan, Ji Li, Yizeng Han, Shiji Song, Gao Huang, Xiu Li |
| 2024 | BootPIG: Bootstrapping Zero-Shot Personalized Image Generation Capabilities in Pretrained Diffusion Models. | Senthil Purushwalkam, Akash Gokul, Shafiq Joty, Nikhil Naik |
| 2024 | Robustness Tokens: Towards Adversarial Robustness of Transformers. | Brian Pulfer, Yury Belousov, Slava Voloshynovskiy |
| 2024 | Sketch & Paint: Stroke-by-Stroke Evolution of Visual Artworks. | Jeripothula Prudviraj, Vikram Jamwal |
| 2024 | Quantization-Friendly Winograd Transformations for Convolutional Neural Networks. | Vladimir Protsenko, Vladimir Kryzhanovskiy, Alexander Filippov |
| 2024 | Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos. | Ekta Prashnani, Koki Nagano, Shalini De Mello, David Luebke, Orazio Gallo |
| 2024 | 3D Hand Pose Estimation in Everyday Egocentric Images. | Aditya Prakash, Ruisen Tu, Matthew Chang, Saurabh Gupta |
| 2024 | Mitigating Perspective Distortion-Induced Shape Ambiguity in Image Crops. | Aditya Prakash, Arjun Gupta, Saurabh Gupta |
| 2024 | 3D Reconstruction of Objects in Hands Without Real World 3D Supervision. | Aditya Prakash, Matthew Chang, Matthew Jin, Ruisen Tu, Saurabh Gupta |
| 2024 | DySeT: A Dynamic Masked Self-distillation Approach for Robust Trajectory Prediction. | Mozhgan Pourkeshavarz, Junrui Zhang, Amir Rasouli |
| 2024 | ShapeFusion: A 3D Diffusion Model for Localized Shape Editing. | Rolandos Alexandros Potamias, Michail Tarasiou, Stylianos Ploumpis, Stefanos Zafeiriou |
| 2024 | Analysis of Hybrid Compositions in Animation Film with Weakly Supervised Learning. | Mnica Apellaniz Portos, Roberto Labadie Tamayo, Claudius Stemmler, Erwin Feyersinger, Andreas Babic, Franziska Bruckner, Vrth hner, Matthias Zeppelzauer |
| 2024 | Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models. | Samuele Poppi, Tobia Poppi, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara |