Skip to content

David Harwath

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

66

Venues

14

Active years

2013–2026

Best venue rank

A*

Where they publish

Papers

66 indexed papers, newest first.

YearVenueTitleAuthors
2026AAAIMMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence.Sonal Kumar, Simon Sedlcek, Vaibhavi Lokegaonkar, Fernando Lpez, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plicka, Miroslav Hlavcek, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaos, Satish Rahi, Laura Herrera-Alarcn, Satvik Dixit, Rupali S. Patil, Soham Deshmukh, Lasha Koroshinadze, Yao Liu, Leibny Paola Garca-Perera, Eleni Zanou, Themos Stafylakis, Joon Son Chung, David Harwath, Chao Zhang, Dinesh Manocha, Alicia Lozano-Diez, Santosh Kesiraju, Sreyan Ghosh, Ramani Duraiswami
2026ACL[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic.Kwanghee Choi, Eunjung Yeo, Cheol Jun Cho, David Harwath, David R. Mortensen
2026ACLVoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation.Puyuan Peng, Zhisheng Zheng, Shang-Wen Li, Abdelrahman Mohamed, David Harwath
2026ACLLinear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration.Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang, Ming-Hao Hsu, Florian Eichin, Zhizheng Wu, Alane Suhr, Michael A. Hedderich, David Harwath, David R. Mortensen, Barbara Plank
2025ASRUUnifying Model and Layer Fusion for Speech Foundation Models.Yi-Jen Shih, David Harwath
2025ASRUCodec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs.Wei-Cheng Tseng, David Harwath
2025EMNLPScaling Rich Style-Prompted Text-to-Speech Datasets.Anuj Diwan, Zhisheng Zheng, David Harwath, Eunsol Choi
2025EMNLPVoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing.Zhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh, Xiaohang Sun, Zhu Liu, Vimal Bhat, David Harwath
2025ICCVVoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models.Sung-Bin Kim, Jeongsoo Choi, Puyuan Peng, Joon Son Chung, Tae-Hyun Oh, David Harwath
2025ICLRSyllableLM: Learning Coarse Semantic Units for Speech Language Models.Alan Baade, Puyuan Peng, David Harwath
2025InterspeechProbing the Robustness Properties of Neural Speech Codecs.Wei-Cheng Tseng, David Harwath
2025WACVTemporally Streaming Audio-Visual Synchronization for Real-World Videos.Jordan Voas, Wei-Cheng Tseng, Layne Berry, Xixi Hu, Puyuan Peng, James Stuedemann, David Harwath
2024ACLVoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild.Puyuan Peng, Po-Yao Huang, Shang-Wen Li, Abdelrahman Mohamed, David Harwath
2024ACLMultimodal Contextualized Semantic Parsing from Speech.Jordan Voas, David Harwath, Raymond Mooney
2024CVPRSoundingActions: Learning How Actions Sound from Narrated Egocentric Videos.Changan Chen, Kumar Ashutosh, Rohit Girdhar, David Harwath, Kristen Grauman
2024ECCVAction2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos.Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue, Wei-Ning Hsu, David Harwath, Kristen Grauman
2024EMNLPTextless Speech-to-Speech Translation With Limited Parallel Data.Anuj Diwan, Anirudh Srinivasan, David Harwath, Eunsol Choi
2024ICASSPIntegrating Self-Supervised Speech Model with Pseudo Word-Level Targets from Visually-Grounded Speech Model.Hung-Chieh Fang, Nai-Xuan Ye, Yi-Jen Shih, Puyuan Peng, Hsuan-Fu Wang, Layne Berry, Hung-Yi Lee, David Harwath
2024ICASSPAV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models.Yuan Tseng, Layne Berry, Yiting Chen, I-Hsiang Chiu, Hsuan-Hao Lin, Max Liu, Puyuan Peng, Yi-Jen Shih, Hung-Yu Wang, Haibin Wu, Poyao Huang, Chun-Mao Lai, Shang-Wen Li, David Harwath, Yu Tsao, Abdelrahman Mohamed, Chi-Luen Feng, Hung-Yi Lee
2024ICASSPSpeechCLIP+: Self-Supervised Multi-Task Representation Learning for Speech Via Clip and Speech-Image Data.Hsuan-Fu Wang, Yi-Jen Shih, Heng-Jui Chang, Layne Berry, Puyuan Peng, Hung-Yi Lee, Hsin-Min Wang, David Harwath
2024ICMLBAT: Learning to Reason about Spatial Sounds with Large Language Models.Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi, David Harwath
2024InterspeechNeural Codec Language Models for Disentangled and Textless Voice Conversion.Alan Baade, Puyuan Peng, David Harwath
2024InterspeechDirect Speech Synthesis from Non-Invasive, Neuromagnetic Signals.Jinuk Kwon, David Harwath, Debadatta Dash, Paul Ferrari, Jun Wang
2024InterspeechImproving Audio Classification with Low-Sampled Microphone Input: An Empirical Study Using Model Self-Distillation.Dawei Liang, Alice Zhang, David Harwath, Edison Thomaz
2024InterspeechInterface Design for Self-Supervised Speech Models.Yi-Jen Shih, David Harwath
2023ACLWhen to Use Efficient Self Attention? Profiling Text, Speech and Image Transformer Variants.Anuj Diwan, Eunsol Choi, David Harwath
2023ASRUAudio-Visual Neural Syntax Acquisition.Cheng-I Jeff Lai, Freda Shi, Puyuan Peng, Yoon Kim, Kevin Gimpel, Shiyu Chang, Yung-Sung Chuang, Saurabhchand Bhati, David D. Cox, David Harwath, Yang Zhang, Karen Livescu, James R. Glass
2023ICASSPM-SpeechCLIP: Leveraging Large-Scale, Pre-Trained Models for Multilingual Speech to Image Retrieval.Layne Berry, Yi-Jen Shih, Hsuan-Fu Wang, Heng-Jui Chang, Hung-Yi Lee, David Harwath
2023ICASSPLearning Audio-Visual Dereverberation.Changan Chen, Wei Sun, David Harwath, Kristen Grauman
2023ICASSPContinual Learning for On-Device Speech Recognition Using Disentangled Conformers.Anuj Diwan, Ching-Feng Yeh, Wei-Ning Hsu, Paden Tomasello, Eunsol Choi, David Harwath, Abdelrahman Mohamed
2023ICASSPUnsupervised Fine-Tuning Data Selection for ASR Using Self-Supervised Speech Models.Reem Gody, David Harwath
2023ICASSPA Dataset for Foreground Speech Analysis With Smartwatches In Everyday Home Environments.Dawei Liang, Zifan Xu, Yinuo Chen, Rebecca Adaimi, David Harwath, Edison Thomaz
2023ICASSPC2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval.Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas, Rogrio Feris, Brian Kingsbury, Leonid Karlinsky, David Harwath, Hilde Kuehne, James R. Glass
2023ICLRContrastive Audio-Visual Masked Autoencoder.Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, James R. Glass
2023InterspeechStyle-transfer based Speech and Audio-visual Scene understanding for Robot Action Sequence Acquisition from Videos.Chiori Hori, Puyuan Peng, David Harwath, Xinyu Liu, Kei Ota, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux
2023InterspeechSyllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model.Puyuan Peng, Shang-Wen Li, Okko Rsnen, Abdelrahman Mohamed, David Harwath
2023InterspeechPrompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task Generalization.Puyuan Peng, Brian Yan, Shinji Watanabe, David Harwath
2023InterspeechComparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages.Andrew Rouditchenko, Sameer Khurana, Samuel Thomas, Rogrio Feris, Leonid Karlinsky, Hilde Kuehne, David Harwath, Brian Kingsbury, James R. Glass
2023IROSLearning to Map Efficiently by Active Echolocation.Xixi Hu, Senthil Purushwalkam, David Harwath, Kristen Grauman
2022CVPREverything at Once - Multi-modal Fusion Transformer for Video Retrieval.Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Hilde Kuehne
2022EMNLPWhy is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality.Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, Kyle Mahowald
2022ICASSPFast-Slow Transformer for Visually Grounding Speech.Puyuan Peng, David Harwath
2022ICASSPAdversarial Input Ablation for Audio-Visual Learning.David Xu, David Harwath
2022InterspeechMAE-AST: Masked Autoencoding Audio Spectrogram Transformer.Alan Baade, Puyuan Peng, David Harwath
2022InterspeechExploring Few-Shot Fine-Tuning Strategies for Models of Visually Grounded Speech.Tyler Miller, David Harwath
2022InterspeechWord Discovery in Visually Grounded, Self-Supervised Speech Models.Puyuan Peng, David Harwath
2022LRECSpeak: A Toolkit Using Amazon Mechanical Turk to Collect and Validate Speech Audio Recordings.Christopher Song, David Harwath, Tuka Alhanai, James R. Glass
2021ACLText-Free Image-to-Speech Synthesis Using Learned Segmental Units.Wei-Ning Hsu, David Harwath, Tyler Miller, Christopher Song, James R. Glass
2021CVPRSpoken Moments: Learning Joint Audio-Visual Representations From Video Descriptions.Mathew Monfort, SouYoung Jin, Alexander H. Liu, David Harwath, Rogrio Feris, James R. Glass, Aude Oliva
2021ICCVMultimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos.Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang
2021InterspeechCascaded Multilingual Audio-Visual Learning from Videos.Andrew Rouditchenko, Angie W. Boggust, David Harwath, Samuel Thomas, Hilde Kuehne, Brian Chen, Rameswar Panda, Rogrio Feris, Brian Kingsbury, Michael Picheny, James R. Glass
2021InterspeechAVLnet: Learning Audio-Visual Language Representations from Instructional Videos.Andrew Rouditchenko, Angie W. Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogrio Schmidt Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, James R. Glass
2020ICASSPTrilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms.Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David Harwath, James R. Glass
2020ICLRLearning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech.David Harwath, Wei-Ning Hsu, James R. Glass
2020InterspeechPair Expansion for Learning Multilingual Semantic Embeddings Using Disjoint Visually-Grounded Speech Audio Datasets.Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David Harwath, James R. Glass
2019CVPRGrounding Spoken Words in Unlabeled Video.Angie W. Boggust, Kartik Audhkhasi, Dhiraj Joshi, David Harwath, Samuel Thomas, Rogrio Schmidt Feris, Danny Gutfreund, Yang Zhang, Antonio Torralba, Michael Picheny, James R. Glass
2019CVPRLearning Words by Drawing Images.Didac Suris, Adri Recasens, David Bau, David Harwath, James R. Glass, Antonio Torralba
2019ICASSPTowards Visually Grounded Sub-word Speech Unit Discovery.David Harwath, James R. Glass
2019InterspeechTowards Bilingual Lexicon Discovery From Visually Grounded Speech Audio.Emmanuel Azuh, David Harwath, James R. Glass
2019InterspeechTransfer Learning from Audio-Visual Grounding to Speech Recognition.Wei-Ning Hsu, David Harwath, James R. Glass
2018ECCVJointly Discovering Visual Objects and Spoken Words from Raw Sensory Input.David Harwath, Adri Recasens, Ddac Surs, Galen Chuang, Antonio Torralba, James R. Glass
2018ICASSPVision as an Interlingua: Learning Multilingual Semantic Embeddings of Untranscribed Speech.David Harwath, Galen Chuang, James R. Glass
2017ACLLearning Word-Like Units from Joint Audio-Visual Analysis.David Harwath, James R. Glass
2017ASRULearning modality-invariant representations for speech and images.Kenneth Leidal, David Harwath, James R. Glass
2014InterspeechChoosing useful word alternates for automatic speech recognition correction interfaces.David Harwath, Alexander Gruenstein, Ian McGraw
2013ICASSPA summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition.Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson, Sanjeev Khudanpur, Kenneth Church, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin Bennett, Benjamin Brschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas