Skip to content

Style-transfer based Speech and Audio-visual Scene understanding for Robot Action Sequence Acquisition from Videos.

Chiori Hori, Puyuan Peng, David Harwath, Xinyu Liu, Kei Ota, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux

Year2023
ProceedingsINTERSPEECH

Browse the full Interspeech paper archive.