| 2026 | ACL | LaMI: Augmenting Large Language Models via Late Multi-Image Fusion. | Guy Yariv, Idan Schwartz, Yossi Adi, Sagie Benaim |
| 2026 | WACV | Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models. | Oz Zafar, Yuval Cohen, Lior Wolf, Idan Schwartz |
| 2024 | AAAI | Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation. | Guy Yariv, Itai Gat, Sagie Benaim, Lior Wolf, Idan Schwartz, Yossi Adi |
| 2023 | BMVC | Zero-Shot Video Captioning by Evolving Pseudo-tokens. | Yoad Tewel, Yoav Shalev, Roy Nadler, Idan Schwartz, Lior Wolf |
| 2023 | ICCV | Discriminative Class Tokens for Text-to-Image Diffusion Models. | Idan Schwartz, Vsteinn Snbjarnarson, Hila Chefer, Serge J. Belongie, Lior Wolf, Sagie Benaim |
| 2023 | Interspeech | Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation. | Guy Yariv, Itai Gat, Lior Wolf, Yossi Adi, Idan Schwartz |
| 2022 | AAAI | Latent Space Explanation by Intervention. | Itai Gat, Guy Lorberbom, Idan Schwartz, Tamir Hazan |
| 2022 | CVPR | ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic. | Yoad Tewel, Yoav Shalev, Idan Schwartz, Lior Wolf |
| 2022 | EMNLP | Describing Sets of Images with Textual-PCA. | Oded Hupert, Idan Schwartz, Lior Wolf |
| 2022 | WACV | Video and Text Matching with Conditioned Embeddings. | Ameen Ali, Idan Schwartz, Tamir Hazan, Lior Wolf |
| 2021 | NAACL | Ensemble of MRR and NDCG models for Visual Dialog. | Idan Schwartz |
| 2019 | CVPR | A Simple Baseline for Audio-Visual Scene-Aware Dialog. | Idan Schwartz, Alexander G. Schwing, Tamir Hazan |
| 2019 | CVPR | Factor Graph Attention. | Idan Schwartz, Seunghak Yu, Tamir Hazan, Alexander G. Schwing |