Skip to content

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation.

Weiting Tan, Jiachen Lian, Hirofumi Inaguma, Paden Tomasello, Philipp Koehn, Xutai Ma

VenueA*EMNLP
Year2025
ProceedingsEMNLP (Findings)

Browse the full EMNLP paper archive.