A Transformer-Based Audio Captioning Model with Keyword Estimation.
Yuma Koizumi, Ryo Masumura, Kyosuke Nishida, Masahiro Yasuda, Shoichiro Saito
Browse the full Interspeech paper archive.
Yuma Koizumi, Ryo Masumura, Kyosuke Nishida, Masahiro Yasuda, Shoichiro Saito
Browse the full Interspeech paper archive.