Skip to content

Michael S. Ryoo

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

100

Venues

23

Active years

2005–2026

Best venue rank

A*

Where they publish

Papers

100 indexed papers, newest first.

YearVenueTitleAuthors
2026EACLToo Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA.Jongwoo Park, Kanchana Ranasinghe, Kumara Kahatapitiya, Wonjeong Ryu, Donghyun Kim, Michael S. Ryoo
2025ACLLAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback.Thai Quoc Hoang, Kung-Hsiang Huang, Shirley Kokane, Jianguo Zhang, Zuxin Liu, Ming Zhu, Jake Grigsby, Tian Lan, Michael S. Ryoo, Chien-Sheng Wu, Shelby Heinecke, Huan Wang, Silvio Savarese, Caiming Xiong, Juan Carlos Niebles
2025ACLLanguage Repository for Long Video Understanding.Kumara Kahatapitiya, Kanchana Ranasinghe, Jongwoo Park, Michael S. Ryoo
2025CVPRGo-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise.Ryan D. Burgert, Yuancheng Xu, Wenqi Xian, Oliver Pilarski, Pascal Clausen, Mingming He, Li Ma, Yitong Deng, Lingxiao Li, Mohsen Mousavi, Michael S. Ryoo, Paul E. Debevec, Ning Yu
2025ICCVAdaptive Caching for Faster Video Generation With Diffusion Transformers.Kumara Kahatapitiya, Haozhe Liu, Sen He, Ding Liu, Menglin Jia, Chenyang Zhang, Michael S. Ryoo, Tian Xie
2025ICCVBLIP-3: A Family of Open Large Multimodal Models.Le Xue, Manli Shu, Anas Awadalla, Jun Wang, An Yan, Senthil Purushwalkam, Honglu Zhou, Viraj Prabhu, Yutong Dai, Michael S. Ryoo, Shrikant Kendre, Jieyu Zhang, Shao-Yen Tseng, Gustavo A. Lujan-Moreno, Matthew L. Olson, Musashi Hinck, David Cobbley, Vasudev Lal, Can Qin, Shu Zhang, Chia-Chih Chen, Ning Yu, Juntao Tan, Tulika Manoj Awalgaonkar, Shelby Heinecke, Huan Wang, Yejin Choi, Ludwig Schmidt, Zeyuan Chen, Silvio Savarese, Juan Carlos Niebles, Caiming Xiong, Ran Xu
2025ICCVStrefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data.Honglu Zhou, Xiangyu Peng, Shrikant Kendre, Michael S. Ryoo, Silvio Savarese, Caiming Xiong, Juan Carlos Niebles
2025ICLRLLaRA: Supercharging Robot Learning Data for Vision-Language Policy.Xiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan D. Burgert, Mu Cai, Yong Jae Lee, Michael S. Ryoo
2025ICLRUnderstanding Long Videos with Multimodal Language Models.Kanchana Ranasinghe, Xiang Li, Kumara Kahatapitiya, Michael S. Ryoo
2024CVPRMAGICK: A Large-Scale Captioned Dataset from Matting Generated Images Using Chroma Keying.Ryan D. Burgert, Brian L. Price, Jason Kuen, Yijun Li, Michael S. Ryoo
2024CVPRVicTR: Video-conditioned Text Representations for Activity Recognition.Kumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, Michael S. Ryoo
2024CVPRMirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual Modalities.A. J. Piergiovanni, Isaac Noble, Dahun Kim, Michael S. Ryoo, Victor Gomes, Anelia Angelova
2024CVPRLearning to Localize Objects Improves Spatial Reasoning in Visual-LLMs.Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo, Tsung-Yu Lin
2024ECCVCoPT: Unsupervised Domain Adaptive Segmentation Using Domain-Agnostic Text Embeddings.Cristina Mata, Kanchana Ranasinghe, Michael S. Ryoo
2024ECCVImage Translation with Kernel Prediction Networks for Semantic Segmentation.Cristina Mata, Michael S. Ryoo, Henrik Turbell
2024ECCVxGen-VideoSyn-1: High-Fidelity Text-to-Video Synthesis with Compressed Representations.Can Qin, Congying Xia, Krithika Ramakrishnan, Michael S. Ryoo, Lifu Tu, Yihao Feng, Manli Shu, Honglu Zhou, Anas Awadalla, Jun Wang, Senthil Purushwalkam, Le Xue, Yingbo Zhou, Huan Wang, Silvio Savarese, Juan Carlos Niebles, Zeyuan Chen, Ran Xu, Caiming Xiong
2024ICRASARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust Attention.Isabel Leal, Krzysztof Choromanski, Deepali Jain, Avinava Dubey, Jake Varley, Michael S. Ryoo, Yao Lu, Frederick Liu, Vikas Sindhwani, Quan Vuong, Tams Sarls, Ken Oslund, Karol Hausman, Kanishka Rao
2024ICRACrossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning.Xiang Li, Varun Belagali, Jinghuan Shang, Michael S. Ryoo
2024SIGGRAPHDiffusion Illusions: Hiding Images in Plain Sight.Ryan D. Burgert, Xiang Li, Abe Leite, Kanchana Ranasinghe, Michael S. Ryoo
2024WACVGrafting Vision Transformers.Jongwoo Park, Kumara Kahatapitiya, Donghyun Kim, Shivchander Sudalairaj, Quanfu Fan, Michael S. Ryoo
2024WACVLimited Data, Unlimited Potential: A Study on ViTs Augmented by Masked Autoencoders.Srijan Das, Tanmay Jain, Dominick Reilly, Pranav Balaji, Soumyajit Karmakar, Shyam Marjit, Xiang Li, Abhijit Das, Michael S. Ryoo
2023AAAIWeakly-Guided Self-Supervised Pretraining for Temporal Activity Detection.Kumara Kahatapitiya, Zhou Ren, Haoxiang Li, Zhenyu Wu, Michael S. Ryoo, Gang Hua
2023BMVCAttributes-Aware Network for Temporal Action Detection.Rui Dai, Srijan Das, Michael S. Ryoo, Franois Brmond
2023CoRLRT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, Quan Vuong, Vincent Vanhoucke, Huong T. Tran, Radu Soricut, Anikait Singh, Jaspiar Singh, Pierre Sermanet, Pannag R. Sanketi, Grecia Salazar, Michael S. Ryoo, Krista Reymann, Kanishka Rao, Karl Pertsch, Igor Mordatch, Henryk Michalewski, Yao Lu, Sergey Levine, Lisa Lee, Tsang-Wei Edward Lee, Isabel Leal, Yuheng Kuang, Dmitry Kalashnikov, Ryan Julian, Nikhil J. Joshi, Alex Irpan, Brian Ichter, Jasmine Hsu, Alexander Herzog, Karol Hausman, Keerthana Gopalakrishnan, Chuyuan Fu, Pete Florence, Chelsea Finn, Kumar Avinava Dubey, Danny Driess, Tianli Ding, Krzysztof Marcin Choromanski, Xi Chen, Yevgen Chebotar, Justice Carbajal, Noah Brown, Anthony Brohan, Montserrat Gonzalez Arenas, Kehang Han
2023CVPRToken Turing Machines.Michael S. Ryoo, Keerthana Gopalakrishnan, Kumara Kahatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu, Julian Ibarz, Anurag Arnab
2023ICLRSocratic Models: Composing Zero-Shot Multimodal Reasoning with Language.Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Marcin Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael S. Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, Pete Florence
2023IJCAISWAT: Spatial Structure Within and Among Tokens.Kumara Kahatapitiya, Michael S. Ryoo
2023ICRAOpen-vocabulary Queryable Scene Representations for Real World Planning.Boyuan Chen, Fei Xia, Brian Ichter, Kanishka Rao, Keerthana Gopalakrishnan, Michael S. Ryoo, Austin Stone, Daniel Kappler
2023ICRAEnergy-Based Models for Cross-Modal Localization using Convolutional Transformers.Alan Wu, Michael S. Ryoo
2023MVACross-modal Manifold Cutmix for Self-supervised Video Representation Learning.Srijan Das, Michael S. Ryoo
2023WACVViewCLR: Learning Self-supervised Video Representation for Unseen Viewpoints.Srijan Das, Michael S. Ryoo
2022CoRLTRITON: Neural Neural Textures for Better Sim2Real.Ryan D. Burgert, Jinghuan Shang, Xiang Li, Michael S. Ryoo
2022CVPRMS-TCT: Multi-Scale Temporal ConvTransformer for Action Detection.Rui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo, Franois Brmond
2022CVPRSelf-supervised Video Transformer.Kanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan, Michael S. Ryoo
2022ECCVVideo Question Answering with Iterative Video-Text Co-tokenization.A. J. Piergiovanni, Kairo Morton, Weicheng Kuo, Michael S. Ryoo, Anelia Angelova
2022ECCVStARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning.Jinghuan Shang, Kumara Kahatapitiya, Xiang Li, Michael S. Ryoo
2022ICLRHybrid Random Features.Krzysztof Marcin Choromanski, Han Lin, Haoxian Chen, Arijit Sehanobish, Yuanzhe Ma, Deepali Jain, Jake Varley, Andy Zeng, Michael S. Ryoo, Valerii Likhosherstov, Dmitry Kalashnikov, Vikas Sindhwani, Adrian Weller
2021BMVCUnsupervised Discovery of Actions in Instructional Videos.A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo, Irfan Essa
2021CVPRCoarse-Fine Networks for Temporal Activity Detection in Videos.Kumara Kahatapitiya, Michael S. Ryoo
2021CVPRAdaptive Intermediate Representations for Video Understanding.Juhana Kangaspunta, A. J. Piergiovanni, Rico Jonschkowski, Michael S. Ryoo, Anelia Angelova
2021CVPRRecognizing Actions in Videos From Unseen Viewpoints.A. J. Piergiovanni, Michael S. Ryoo
2021ICCV4D-Net for Learned Multi-Modal Alignment.A. J. Piergiovanni, Vincent Casser, Michael S. Ryoo, Anelia Angelova
2021ICRAVisionary: Vision architecture discovery for robot learning.Iretiayo Akinola, Anelia Angelova, Yao Lu, Yevgen Chebotar, Dmitry Kalashnikov, Jacob Varley, Julian Ibarz, Michael S. Ryoo
2021IROSSelf-Supervised Disentangled Representation Learning for Third-Person Imitation Learning.Jinghuan Shang, Michael S. Ryoo
2020AAAIDifferentiable Grammars for Videos.A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo
2020CVPREvolving Losses for Unsupervised Video Representation Learning.A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo
2020ECCVPassword-Conditioned Anonymization and Deanonymization with Face Identity Transformers.Xiuye Gu, Weixin Luo, Michael S. Ryoo, Yong Jae Lee
2020ECCVAdversarial Generative Grammars for Human Activity Prediction.A. J. Piergiovanni, Anelia Angelova, Alexander Toshev, Michael S. Ryoo
2020ECCVAssembleNet++: Assembling Modality Representations via Attention Connections.Michael S. Ryoo, A. J. Piergiovanni, Juhana Kangaspunta, Anelia Angelova
2020ECCVAttentionNAS: Spatiotemporal Attention Cell Search for Video Classification.Xiaofang Wang, Xuehan Xiong, Maxim Neumann, A. J. Piergiovanni, Michael S. Ryoo, Anelia Angelova, Kris M. Kitani, Wei Hua
2020ICLRAssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures.Michael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia Angelova
2020WACVLearning Multimodal Representations for Unseen Activities.A. J. Piergiovanni, Michael S. Ryoo
2019CoRLModel-based Behavioral Cloning with Future Image Similarity Learning.Alan Wu, A. J. Piergiovanni, Michael S. Ryoo
2019CVPRRepresentation Flow for Action Recognition.A. J. Piergiovanni, Michael S. Ryoo
2019CVPREarly Detection of Injuries in MLB Pitchers From Video.A. J. Piergiovanni, Michael S. Ryoo
2019DACRobustly Executing DNNs in IoT Systems Using Coded Distributed Computing.Ramyad Hadidi, Jiashen Cao, Michael S. Ryoo, Hyesoon Kim
2019ICCVEvolving Space-Time Neural Architectures for Videos.A. J. Piergiovanni, Anelia Angelova, Alexander Toshev, Michael S. Ryoo
2019ICMLTemporal Gaussian Mixture Layer for Videos.A. J. Piergiovanni, Michael S. Ryoo
2019IROSPrivacy-Preserving Robot Vision with Anonymized Faces by Extreme Low Resolution.Myeung Un Kim, Harim Lee, Hyun Jong Yang, Michael S. Ryoo
2019IROSLearning Real-World Robot Policies by Dreaming.A. J. Piergiovanni, Alan Wu, Michael S. Ryoo
2018AAAIExtreme Low Resolution Activity Recognition With Multi-Siamese Embedding Learning.Michael S. Ryoo, Kiyoon Kim, Hyun Jong Yang
2018ASPLOSReal-Time Image Recognition Using Collaborative IoT Devices.Ramyad Hadidi, Jiashen Cao, Matthew Woodward, Michael S. Ryoo, Hyesoon Kim
2018CVPRLearning Latent Super-Events to Detect Multiple Activities in Videos.A. J. Piergiovanni, Michael S. Ryoo
2018CVPRFine-Grained Activity Recognition in Baseball Videos.A. J. Piergiovanni, Michael S. Ryoo
2018CVPRAction-Conditioned Convolutional Future Regression Models for Robot Imitation Learning.Alan Wu, A. J. Piergiovanni, Michael S. Ryoo
2018ECCVForecasting Hands and Objects in Future Frames.Chenyou Fan, Jangwon Lee, Michael S. Ryoo
2018ECCVLearning to Anonymize Faces for Privacy Preserving Action Detection.Zhongzheng Ren, Yong Jae Lee, Michael S. Ryoo
2018ECCVJoint Person Segmentation and Identification in Synchronized First- and Third-Person Videos.Mingze Xu, Chenyou Fan, Yuchen Wang, Michael S. Ryoo, David J. Crandall
2017AAAITitle Learning Latent Subevents in Activity Videos Using Temporal Attention Filters.A. J. Piergiovanni, Chenyou Fan, Michael S. Ryoo
2017AAAIPrivacy-Preserving Human Activity Recognition from Extreme Low Resolution.Michael S. Ryoo, Brandon Rothrock, Charles Fleming, Hyun Jong Yang
2017CVPRIdentifying First-Person Camera Wearers in Third-Person Videos.Chenyou Fan, Jangwon Lee, Mingze Xu, Krishna Kumar Singh, Yong Jae Lee, David J. Crandall, Michael S. Ryoo
2017CVPRLearning Robot Activities from First-Person Human Videos Using Convolutional Future Regression.Jangwon Lee, Michael S. Ryoo
2017IJCAIMulti-Type Activity Recognition from a Robot's Viewpoint.Ilaria Gori, J. K. Aggarwal, Larry H. Matthies, Michael S. Ryoo
2017IROSLearning robot activities from first-person human videos using convolutional future regression.Jangwon Lee, Michael S. Ryoo
2017ICRALearning social affordance grammar from videos: Transferring human interactions to human-robot interactions.Tianmin Shu, Xiaofeng Gao, Michael S. Ryoo, Song-Chun Zhu
2016IJCAILearning Social Affordance for Human-Robot Interaction.Tianmin Shu, Michael S. Ryoo, Song-Chun Zhu
2015CVPRPooled motion features for first-person videos.Michael S. Ryoo, Brandon Rothrock, Larry H. Matthies
2015HRIRobot-Centric Activity Prediction from First-Person Videos: What Will They Do to Me'.Michael S. Ryoo, Thomas J. Fuchs, Lu Xia, Jake K. Aggarwal, Larry H. Matthies
2015WACVRobot-centric Activity Recognition from First-Person RGB-D Videos.Lu Xia, Ilaria Gori, Jake K. Aggarwal, Michael S. Ryoo
2014CVPRAn Introduction to the 3rd Workshop on Egocentric (First-Person) Vision.Steve Mann, Kris M. Kitani, Yong Jae Lee, Michael S. Ryoo, Alireza Fathi
2014ICPRFirst-Person Animal Activity Recognition from Egocentric Videos.Yumi Iwashita, Asamichi Takamine, Ryo Kurazume, Michael S. Ryoo
2013BMVCRecognizing Humans in Motion: Trajectory-based Aerial Video Analysis.Yumi Iwashita, Michael S. Ryoo, Thomas J. Fuchs, Curtis Padgett
2013CVPRFirst-Person Activity Recognition: What Are They Doing to Me?Michael S. Ryoo, Larry H. Matthies
2012IROSReliable object detection and segmentation using inpainting.Ji Hoon Joung, Michael S. Ryoo, Sunglok Choi, Sung-Rak Kim
2011ICCVHuman activity prediction: Early recognition of ongoing activities from streaming videos.Michael S. Ryoo
2011WACVPersonal driving diary: Constructing a video archive of everyday driving events.Michael S. Ryoo, Jae-Yeong Lee, Ji Hoon Joung, Sunglok Choi, Wonpil Yu
2011WACVOne video is sufficient? Human activity recognition using active video composition.Michael S. Ryoo, Wonpil Yu
2010ICPRAn Overview of Contest on Semantic Description of Human Activities (SDHA) 2010.Michael S. Ryoo, Chia-Chih Chen, J. K. Aggarwal, Amit K. Roy-Chowdhury
2009CVPRStochastic representation and recognition of high-level group activities: Describing structural uncertainties in human activities.Michael S. Ryoo, Jake K. Aggarwal
2009ICCVSpatio-temporal relationship match: Video structure comparison for recognition of complex human activities.Michael S. Ryoo, Jake K. Aggarwal
2008CVPRObserve-and-explain: A new approach for multiple hypotheses tracking of humans and objects.Michael S. Ryoo, Jake K. Aggarwal
2008ICPRHuman activities: Handling uncertainties using fuzzy time intervals.Michael S. Ryoo, Jake K. Aggarwal
2007AVSSDetection of abandoned objects in crowded environments.Medha Bhargava, Chia-Chih Chen, Michael S. Ryoo, Jake K. Aggarwal
2007AVSSReal-time detection of illegally parked vehicles using 1-D transformation.Jong Taek Lee, Michael S. Ryoo, Matthew Riley, Jake K. Aggarwal
2007CVPRHierarchical Recognition of Human Activities Interacting with Objects.Michael S. Ryoo, J. K. Aggarwal
2007IJCAIRobust Human-Computer Interaction System Guiding a User by Providing Feedback.Michael S. Ryoo, Jake K. Aggarwal
2006CVPRRecognition of Composite Human Activities through Context-Free Grammar Based Representation.Michael S. Ryoo, J. K. Aggarwal
2006ICPRSemantic Understanding of Continued and Recursive Human Activities.Michael S. Ryoo, J. K. Aggarwal
2005ACIIAffective Dialogue Communication System with Emotional Memories for Humanoid Robots.Michael S. Ryoo, Yongho Seo, Hye-Won Jung, Hyun Seung Yang
2005GECCOEvolving neural network ensembles for control problems.David Pardoe, Michael S. Ryoo, Risto Miikkulainen