Few-Shot Action Scene Graph Generation from Video via Multimodal Language Models for Structuring Spatial Experience.
Jinseok Hong, Hyerim Park, Heejeong Ko, Woontack Woo
Browse the full ISMAR paper archive.
Jinseok Hong, Hyerim Park, Heejeong Ko, Woontack Woo
Browse the full ISMAR paper archive.