Semantic retrieval of personal photos using a deep autoencoder fusing visual features with speech annotations represented as word/paragraph vectors.
Hung-tsung Lu, Yuan-ming Liou, Hung-yi Lee, Lin-Shan Lee
Browse the full Interspeech paper archive.
Hung-tsung Lu, Yuan-ming Liou, Hung-yi Lee, Lin-Shan Lee
Browse the full Interspeech paper archive.