Skip to content

A. J. Piergiovanni

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

32

Venues

11

Active years

2015–2025

Best venue rank

A*

Where they publish

Papers

32 indexed papers, newest first.

YearVenueTitleAuthors
2025CVPRVideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models.Dahun Kim, A. J. Piergiovanni, Ganesh Satish Mallya, Anelia Angelova
2024CVPROn Scaling Up a Multilingual Vision and Language Model.Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Riquelme Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, Siamak Shakeri, Mostafa Dehghani, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Peter Steiner, Yang Li, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut
2024CVPRMirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual Modalities.A. J. Piergiovanni, Isaac Noble, Dahun Kim, Michael S. Ryoo, Victor Gomes, Anelia Angelova
2024WACVSLVP: Self-Supervised Language-Video Pre-Training for Referring Video Object Segmentation.Jie Mei, A. J. Piergiovanni, Jenq-Neng Hwang, Wei Li
2023CVPRRethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning.A. J. Piergiovanni, Weicheng Kuo, Anelia Angelova
2023ICLRCompound Tokens: Channel Fusion for Vision-Language Representation Learning.Maxwell Mbabilla Aladago, A. J. Piergiovanni
2023ICLRPaLI: A Jointly-Scaled Multilingual Language-Image Model.Xi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish V. Thapliyal, James Bradbury, Weicheng Kuo
2023ICLROpen-Vocabulary Object Detection upon Frozen Vision and Language Models.Weicheng Kuo, Yin Cui, Xiuye Gu, A. J. Piergiovanni, Anelia Angelova
2022ECCVFindIt: Generalized Localization with Natural Language Queries.Weicheng Kuo, Fred Bertsch, Wei Li, A. J. Piergiovanni, Mohammad Saffar, Anelia Angelova
2022ECCVVideo Question Answering with Iterative Video-Text Co-tokenization.A. J. Piergiovanni, Kairo Morton, Weicheng Kuo, Michael S. Ryoo, Anelia Angelova
2021BMVCUnsupervised Discovery of Actions in Instructional Videos.A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo, Irfan Essa
2021CVPRAdaptive Intermediate Representations for Video Understanding.Juhana Kangaspunta, A. J. Piergiovanni, Rico Jonschkowski, Michael S. Ryoo, Anelia Angelova
2021CVPRRecognizing Actions in Videos From Unseen Viewpoints.A. J. Piergiovanni, Michael S. Ryoo
2021ICCV4D-Net for Learned Multi-Modal Alignment.A. J. Piergiovanni, Vincent Casser, Michael S. Ryoo, Anelia Angelova
2020AAAIDifferentiable Grammars for Videos.A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo
2020CVPREvolving Losses for Unsupervised Video Representation Learning.A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo
2020ECCVAdversarial Generative Grammars for Human Activity Prediction.A. J. Piergiovanni, Anelia Angelova, Alexander Toshev, Michael S. Ryoo
2020ECCVAssembleNet++: Assembling Modality Representations via Attention Connections.Michael S. Ryoo, A. J. Piergiovanni, Juhana Kangaspunta, Anelia Angelova
2020ECCVAttentionNAS: Spatiotemporal Attention Cell Search for Video Classification.Xiaofang Wang, Xuehan Xiong, Maxim Neumann, A. J. Piergiovanni, Michael S. Ryoo, Anelia Angelova, Kris M. Kitani, Wei Hua
2020ICLRAssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures.Michael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia Angelova
2020WACVLearning Multimodal Representations for Unseen Activities.A. J. Piergiovanni, Michael S. Ryoo
2019CoRLModel-based Behavioral Cloning with Future Image Similarity Learning.Alan Wu, A. J. Piergiovanni, Michael S. Ryoo
2019CVPRRepresentation Flow for Action Recognition.A. J. Piergiovanni, Michael S. Ryoo
2019CVPREarly Detection of Injuries in MLB Pitchers From Video.A. J. Piergiovanni, Michael S. Ryoo
2019ICCVEvolving Space-Time Neural Architectures for Videos.A. J. Piergiovanni, Anelia Angelova, Alexander Toshev, Michael S. Ryoo
2019ICMLTemporal Gaussian Mixture Layer for Videos.A. J. Piergiovanni, Michael S. Ryoo
2019IROSLearning Real-World Robot Policies by Dreaming.A. J. Piergiovanni, Alan Wu, Michael S. Ryoo
2018CVPRLearning Latent Super-Events to Detect Multiple Activities in Videos.A. J. Piergiovanni, Michael S. Ryoo
2018CVPRFine-Grained Activity Recognition in Baseball Videos.A. J. Piergiovanni, Michael S. Ryoo
2018CVPRAction-Conditioned Convolutional Future Regression Models for Robot Imitation Learning.Alan Wu, A. J. Piergiovanni, Michael S. Ryoo
2017AAAITitle Learning Latent Subevents in Activity Videos Using Temporal Attention Filters.A. J. Piergiovanni, Chenyou Fan, Michael S. Ryoo
2015CogSciComputational principles underlying people's behavior explanations.A. J. Piergiovanni, Alan Jern