Visual Latent Captioning - Towards Verbalizing Vision Transformer Encoders.
Sogol Haghighat, Tim Daniel Metzler, Santosh Thoduka, Sebastian Houben
Browse the full ECIR paper archive.
Sogol Haghighat, Tim Daniel Metzler, Santosh Thoduka, Sebastian Houben
Browse the full ECIR paper archive.