Show, Think, and Tell: Thought-Augmented Fine-Tuning of Large Language Models for Video Captioning.
Byoungjip Kim, Dasol Hwang, Sungjun Cho, Youngsoo Jang, Honglak Lee, Moontae Lee
Browse the full CVPR paper archive.
Byoungjip Kim, Dasol Hwang, Sungjun Cho, Youngsoo Jang, Honglak Lee, Moontae Lee
Browse the full CVPR paper archive.