David Harwath
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
66
Venues
14
Active years
2013–2026
Best venue rank
A*
Where they publish
Papers
66 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2026 | AAAI | MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence. | Sonal Kumar, Simon Sedlcek, Vaibhavi Lokegaonkar, Fernando Lpez, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plicka, Miroslav Hlavcek, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaos, Satish Rahi, Laura Herrera-Alarcn, Satvik Dixit, Rupali S. Patil, Soham Deshmukh, Lasha Koroshinadze, Yao Liu, Leibny Paola Garca-Perera, Eleni Zanou, Themos Stafylakis, Joon Son Chung, David Harwath, Chao Zhang, Dinesh Manocha, Alicia Lozano-Diez, Santosh Kesiraju, Sreyan Ghosh, Ramani Duraiswami |
| 2026 | ACL | [b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic. | Kwanghee Choi, Eunjung Yeo, Cheol Jun Cho, David Harwath, David R. Mortensen |
| 2026 | ACL | VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation. | Puyuan Peng, Zhisheng Zheng, Shang-Wen Li, Abdelrahman Mohamed, David Harwath |
| 2026 | ACL | Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration. | Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang, Ming-Hao Hsu, Florian Eichin, Zhizheng Wu, Alane Suhr, Michael A. Hedderich, David Harwath, David R. Mortensen, Barbara Plank |
| 2025 | ASRU | Unifying Model and Layer Fusion for Speech Foundation Models. | Yi-Jen Shih, David Harwath |
| 2025 | ASRU | Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs. | Wei-Cheng Tseng, David Harwath |
| 2025 | EMNLP | Scaling Rich Style-Prompted Text-to-Speech Datasets. | Anuj Diwan, Zhisheng Zheng, David Harwath, Eunsol Choi |
| 2025 | EMNLP | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing. | Zhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh, Xiaohang Sun, Zhu Liu, Vimal Bhat, David Harwath |
| 2025 | ICCV | VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models. | Sung-Bin Kim, Jeongsoo Choi, Puyuan Peng, Joon Son Chung, Tae-Hyun Oh, David Harwath |
| 2025 | ICLR | SyllableLM: Learning Coarse Semantic Units for Speech Language Models. | Alan Baade, Puyuan Peng, David Harwath |
| 2025 | Interspeech | Probing the Robustness Properties of Neural Speech Codecs. | Wei-Cheng Tseng, David Harwath |
| 2025 | WACV | Temporally Streaming Audio-Visual Synchronization for Real-World Videos. | Jordan Voas, Wei-Cheng Tseng, Layne Berry, Xixi Hu, Puyuan Peng, James Stuedemann, David Harwath |
| 2024 | ACL | VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild. | Puyuan Peng, Po-Yao Huang, Shang-Wen Li, Abdelrahman Mohamed, David Harwath |
| 2024 | ACL | Multimodal Contextualized Semantic Parsing from Speech. | Jordan Voas, David Harwath, Raymond Mooney |
| 2024 | CVPR | SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos. | Changan Chen, Kumar Ashutosh, Rohit Girdhar, David Harwath, Kristen Grauman |
| 2024 | ECCV | Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos. | Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue, Wei-Ning Hsu, David Harwath, Kristen Grauman |
| 2024 | EMNLP | Textless Speech-to-Speech Translation With Limited Parallel Data. | Anuj Diwan, Anirudh Srinivasan, David Harwath, Eunsol Choi |
| 2024 | ICASSP | Integrating Self-Supervised Speech Model with Pseudo Word-Level Targets from Visually-Grounded Speech Model. | Hung-Chieh Fang, Nai-Xuan Ye, Yi-Jen Shih, Puyuan Peng, Hsuan-Fu Wang, Layne Berry, Hung-Yi Lee, David Harwath |
| 2024 | ICASSP | AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models. | Yuan Tseng, Layne Berry, Yiting Chen, I-Hsiang Chiu, Hsuan-Hao Lin, Max Liu, Puyuan Peng, Yi-Jen Shih, Hung-Yu Wang, Haibin Wu, Poyao Huang, Chun-Mao Lai, Shang-Wen Li, David Harwath, Yu Tsao, Abdelrahman Mohamed, Chi-Luen Feng, Hung-Yi Lee |
| 2024 | ICASSP | SpeechCLIP+: Self-Supervised Multi-Task Representation Learning for Speech Via Clip and Speech-Image Data. | Hsuan-Fu Wang, Yi-Jen Shih, Heng-Jui Chang, Layne Berry, Puyuan Peng, Hung-Yi Lee, Hsin-Min Wang, David Harwath |
| 2024 | ICML | BAT: Learning to Reason about Spatial Sounds with Large Language Models. | Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi, David Harwath |
| 2024 | Interspeech | Neural Codec Language Models for Disentangled and Textless Voice Conversion. | Alan Baade, Puyuan Peng, David Harwath |
| 2024 | Interspeech | Direct Speech Synthesis from Non-Invasive, Neuromagnetic Signals. | Jinuk Kwon, David Harwath, Debadatta Dash, Paul Ferrari, Jun Wang |
| 2024 | Interspeech | Improving Audio Classification with Low-Sampled Microphone Input: An Empirical Study Using Model Self-Distillation. | Dawei Liang, Alice Zhang, David Harwath, Edison Thomaz |
| 2024 | Interspeech | Interface Design for Self-Supervised Speech Models. | Yi-Jen Shih, David Harwath |
| 2023 | ACL | When to Use Efficient Self Attention? Profiling Text, Speech and Image Transformer Variants. | Anuj Diwan, Eunsol Choi, David Harwath |
| 2023 | ASRU | Audio-Visual Neural Syntax Acquisition. | Cheng-I Jeff Lai, Freda Shi, Puyuan Peng, Yoon Kim, Kevin Gimpel, Shiyu Chang, Yung-Sung Chuang, Saurabhchand Bhati, David D. Cox, David Harwath, Yang Zhang, Karen Livescu, James R. Glass |
| 2023 | ICASSP | M-SpeechCLIP: Leveraging Large-Scale, Pre-Trained Models for Multilingual Speech to Image Retrieval. | Layne Berry, Yi-Jen Shih, Hsuan-Fu Wang, Heng-Jui Chang, Hung-Yi Lee, David Harwath |
| 2023 | ICASSP | Learning Audio-Visual Dereverberation. | Changan Chen, Wei Sun, David Harwath, Kristen Grauman |
| 2023 | ICASSP | Continual Learning for On-Device Speech Recognition Using Disentangled Conformers. | Anuj Diwan, Ching-Feng Yeh, Wei-Ning Hsu, Paden Tomasello, Eunsol Choi, David Harwath, Abdelrahman Mohamed |
| 2023 | ICASSP | Unsupervised Fine-Tuning Data Selection for ASR Using Self-Supervised Speech Models. | Reem Gody, David Harwath |
| 2023 | ICASSP | A Dataset for Foreground Speech Analysis With Smartwatches In Everyday Home Environments. | Dawei Liang, Zifan Xu, Yinuo Chen, Rebecca Adaimi, David Harwath, Edison Thomaz |
| 2023 | ICASSP | C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval. | Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas, Rogrio Feris, Brian Kingsbury, Leonid Karlinsky, David Harwath, Hilde Kuehne, James R. Glass |
| 2023 | ICLR | Contrastive Audio-Visual Masked Autoencoder. | Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, James R. Glass |
| 2023 | Interspeech | Style-transfer based Speech and Audio-visual Scene understanding for Robot Action Sequence Acquisition from Videos. | Chiori Hori, Puyuan Peng, David Harwath, Xinyu Liu, Kei Ota, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux |
| 2023 | Interspeech | Syllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model. | Puyuan Peng, Shang-Wen Li, Okko Rsnen, Abdelrahman Mohamed, David Harwath |
| 2023 | Interspeech | Prompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task Generalization. | Puyuan Peng, Brian Yan, Shinji Watanabe, David Harwath |
| 2023 | Interspeech | Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages. | Andrew Rouditchenko, Sameer Khurana, Samuel Thomas, Rogrio Feris, Leonid Karlinsky, Hilde Kuehne, David Harwath, Brian Kingsbury, James R. Glass |
| 2023 | IROS | Learning to Map Efficiently by Active Echolocation. | Xixi Hu, Senthil Purushwalkam, David Harwath, Kristen Grauman |
| 2022 | CVPR | Everything at Once - Multi-modal Fusion Transformer for Video Retrieval. | Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Hilde Kuehne |
| 2022 | EMNLP | Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality. | Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, Kyle Mahowald |
| 2022 | ICASSP | Fast-Slow Transformer for Visually Grounding Speech. | Puyuan Peng, David Harwath |
| 2022 | ICASSP | Adversarial Input Ablation for Audio-Visual Learning. | David Xu, David Harwath |
| 2022 | Interspeech | MAE-AST: Masked Autoencoding Audio Spectrogram Transformer. | Alan Baade, Puyuan Peng, David Harwath |
| 2022 | Interspeech | Exploring Few-Shot Fine-Tuning Strategies for Models of Visually Grounded Speech. | Tyler Miller, David Harwath |
| 2022 | Interspeech | Word Discovery in Visually Grounded, Self-Supervised Speech Models. | Puyuan Peng, David Harwath |
| 2022 | LREC | Speak: A Toolkit Using Amazon Mechanical Turk to Collect and Validate Speech Audio Recordings. | Christopher Song, David Harwath, Tuka Alhanai, James R. Glass |
| 2021 | ACL | Text-Free Image-to-Speech Synthesis Using Learned Segmental Units. | Wei-Ning Hsu, David Harwath, Tyler Miller, Christopher Song, James R. Glass |
| 2021 | CVPR | Spoken Moments: Learning Joint Audio-Visual Representations From Video Descriptions. | Mathew Monfort, SouYoung Jin, Alexander H. Liu, David Harwath, Rogrio Feris, James R. Glass, Aude Oliva |
| 2021 | ICCV | Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos. | Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang |
| 2021 | Interspeech | Cascaded Multilingual Audio-Visual Learning from Videos. | Andrew Rouditchenko, Angie W. Boggust, David Harwath, Samuel Thomas, Hilde Kuehne, Brian Chen, Rameswar Panda, Rogrio Feris, Brian Kingsbury, Michael Picheny, James R. Glass |
| 2021 | Interspeech | AVLnet: Learning Audio-Visual Language Representations from Instructional Videos. | Andrew Rouditchenko, Angie W. Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogrio Schmidt Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, James R. Glass |
| 2020 | ICASSP | Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms. | Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David Harwath, James R. Glass |
| 2020 | ICLR | Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech. | David Harwath, Wei-Ning Hsu, James R. Glass |
| 2020 | Interspeech | Pair Expansion for Learning Multilingual Semantic Embeddings Using Disjoint Visually-Grounded Speech Audio Datasets. | Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David Harwath, James R. Glass |
| 2019 | CVPR | Grounding Spoken Words in Unlabeled Video. | Angie W. Boggust, Kartik Audhkhasi, Dhiraj Joshi, David Harwath, Samuel Thomas, Rogrio Schmidt Feris, Danny Gutfreund, Yang Zhang, Antonio Torralba, Michael Picheny, James R. Glass |
| 2019 | CVPR | Learning Words by Drawing Images. | Didac Suris, Adri Recasens, David Bau, David Harwath, James R. Glass, Antonio Torralba |
| 2019 | ICASSP | Towards Visually Grounded Sub-word Speech Unit Discovery. | David Harwath, James R. Glass |
| 2019 | Interspeech | Towards Bilingual Lexicon Discovery From Visually Grounded Speech Audio. | Emmanuel Azuh, David Harwath, James R. Glass |
| 2019 | Interspeech | Transfer Learning from Audio-Visual Grounding to Speech Recognition. | Wei-Ning Hsu, David Harwath, James R. Glass |
| 2018 | ECCV | Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input. | David Harwath, Adri Recasens, Ddac Surs, Galen Chuang, Antonio Torralba, James R. Glass |
| 2018 | ICASSP | Vision as an Interlingua: Learning Multilingual Semantic Embeddings of Untranscribed Speech. | David Harwath, Galen Chuang, James R. Glass |
| 2017 | ACL | Learning Word-Like Units from Joint Audio-Visual Analysis. | David Harwath, James R. Glass |
| 2017 | ASRU | Learning modality-invariant representations for speech and images. | Kenneth Leidal, David Harwath, James R. Glass |
| 2014 | Interspeech | Choosing useful word alternates for automatic speech recognition correction interfaces. | David Harwath, Alexander Gruenstein, Ian McGraw |
| 2013 | ICASSP | A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition. | Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson, Sanjeev Khudanpur, Kenneth Church, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin Bennett, Benjamin Brschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas |