| 2025 | ICASSP | Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement. | Jae-Sung Bae, Anastasia Kuznetsova, Dinesh Manocha, John R. Hershey, Trausti T. Kristjansson, Minje Kim |
| 2025 | ICASSP | Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables. | Artem Dementyev, Chandan K. A. Reddy, Scott Wisdom, Navin Chatlani, John R. Hershey, Richard F. Lyon |
| 2025 | ICASSP | Generative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation. | Jackie Lin, Georg Gtz, Hermes Sampedro Llopis, Haukur Hafsteinsson, Steinar Gujnsson, Daniel Gert Nielsen, Finnur Pind, Paris Smaragdis, Dinesh Manocha, John R. Hershey, Trausti T. Kristjansson, Minje Kim |
| 2025 | ICLR | I-Con: A Unifying Framework for Representation Learning. | Shaden Naif Alshammari, John R. Hershey, Axel Feldmann, William T. Freeman, Mark Hamilton |
| 2024 | CVPR | Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language. | Mark Hamilton, Andrew Zisserman, John R. Hershey, William T. Freeman |
| 2024 | ICASSP | Unsupervised Multi-Channel Separation And Adaptation. | Cong Han, Kevin W. Wilson, Scott Wisdom, John R. Hershey |
| 2024 | Interspeech | Unsupervised Improved MVDR Beamforming for Sound Enhancement. | Jacob Kealey, John R. Hershey, Franois Grondin |
| 2023 | ICASSP | Audioslots: A Slot-Centric Generative Model For Audio Separation. | Pradyumna Reddy, Scott Wisdom, Klaus Greff, John R. Hershey, Thomas Kipf |
| 2023 | Interspeech | TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition. | Hakan Erdogan, Scott Wisdom, Xuankai Chang, Zaln Borsos, Marco Tagliasacchi, Neil Zeghidour, John R. Hershey |
| 2022 | ECCV | AudioScopeV2: Audio-Visual Attention Architectures for Calibrated Open-Domain On-Screen Sound Separation. | Efthymios Tzinis, Scott Wisdom, Tal Remez, John R. Hershey |
| 2022 | ICASSP | Improving Bird Classification with Unsupervised Sound Separation. | Tom Denton, Scott Wisdom, John R. Hershey |
| 2022 | ICASSP | Adapting Speech Separation to Real-World Meetings using Mixture Invariant Training. | Aswin Sivaraman, Scott Wisdom, Hakan Erdogan, John R. Hershey |
| 2022 | Interspeech | CycleGAN-based Unpaired Speech Dereverberation. | Hannah Muckenhirn, Aleksandr Safin, Hakan Erdogan, Felix de Chaumont Quitry, Marco Tagliasacchi, Scott Wisdom, John R. Hershey |
| 2022 | Interspeech | Distance-Based Sound Separation. | Katharine Patterson, Kevin W. Wilson, Scott Wisdom, John R. Hershey |
| 2021 | ICASSP | End-To-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings. | Soumi Maiti, Hakan Erdogan, Kevin W. Wilson, Scott Wisdom, Shinji Watanabe, John R. Hershey |
| 2021 | ICASSP | Sound Event Detection and Separation: A Benchmark on Desed Synthetic Soundscapes. | Nicolas Turpault, Romain Serizel, Scott Wisdom, Hakan Erdogan, John R. Hershey, Eduardo Fonseca, Prem Seetharaman, Justin Salamon |
| 2021 | ICASSP | What's all the Fuss about Free Universal Sound Separation Data? | Scott Wisdom, Hakan Erdogan, Daniel P. W. Ellis, Romain Serizel, Nicolas Turpault, Eduardo Fonseca, Justin Salamon, Prem Seetharaman, John R. Hershey |
| 2021 | ICLR | Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds. | Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey, Tal Remez, Dan Ellis, John R. Hershey |
| 2021 | Interspeech | Continuous Speech Separation Using Speaker Inventory for Long Recording. | Cong Han, Yi Luo, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, Zhuo Chen |
| 2020 | ICASSP | Improving Universal Sound Separation Using Sound Classification. | Efthymios Tzinis, Scott Wisdom, John R. Hershey, Aren Jansen, Daniel P. W. Ellis |
| 2019 | ICASSP | SDR - Half-baked or Well Done? | Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, John R. Hershey |
| 2019 | ICASSP | The Phasebook: Building Complex Masks via Discrete Representations for Source Separation. | Jonathan Le Roux, Gordon Wichern, Shinji Watanabe, Andy M. Sarroff, John R. Hershey |
| 2019 | ICASSP | Differentiable Consistency Constraints for Improved Deep Speech Enhancement. | Scott Wisdom, John R. Hershey, Kevin W. Wilson, Jeremy Thorpe, Michael Chinen, Brian Patton, Rif A. Saurous |
| 2019 | Interspeech | End-to-End Multilingual Multi-Speaker Speech Recognition. | Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Jonathan Le Roux, John R. Hershey |
| 2019 | Interspeech | VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking. | Quan Wang, Hannah Muckenhirn, Kevin W. Wilson, Prashant Sridhar, Zelin Wu, John R. Hershey, Rif A. Saurous, Ron J. Weiss, Ye Jia, Ignacio Lpez-Moreno |
| 2018 | ACL | A Purely End-to-End System for Multi-speaker Speech Recognition. | Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Jonathan Le Roux, John R. Hershey |
| 2018 | ICASSP | Speaker Adaptation for Multichannel End-to-End Speech Recognition. | Tsubasa Ochiai, Shinji Watanabe, Shigeru Katagiri, Takaaki Hori, John R. Hershey |
| 2018 | ICASSP | An End-to-End Language-Tracking Speech Recognizer for Mixed-Language Speech. | Hiroshi Seki, Shinji Watanabe, Takaaki Hori, Jonathan Le Roux, John R. Hershey |
| 2018 | ICASSP | End-to-End Multi-Speaker Speech Recognition. | Shane Settle, Jonathan Le Roux, Takaaki Hori, Shinji Watanabe, John R. Hershey |
| 2018 | ICASSP | Multi-Channel Deep Clustering: Discriminative Spectral and Spatial Embeddings for Speaker-Independent Speech Separation. | Zhong-Qiu Wang, Jonathan Le Roux, John R. Hershey |
| 2018 | ICASSP | Alternative Objective Functions for Deep Clustering. | Zhong-Qiu Wang, Jonathan Le Roux, John R. Hershey |
| 2018 | Interspeech | End-to-End Speech Separation with Unfolded Iterative Phase Reconstruction. | Zhong-Qiu Wang, Jonathan Le Roux, DeLiang Wang, John R. Hershey |
| 2017 | ACL | Joint CTC/attention decoding for end-to-end speech recognition. | Takaaki Hori, Shinji Watanabe, John R. Hershey |
| 2017 | ASRU | Early and late integration of audio features for automatic video description. | Chiori Hori, Takaaki Hori, Tim K. Marks, John R. Hershey |
| 2017 | ASRU | Multi-level language modeling and decoding for open vocabulary end-to-end speech recognition. | Takaaki Hori, Shinji Watanabe, John R. Hershey |
| 2017 | ASRU | Language independent end-to-end architecture for joint language identification and speech recognition. | Shinji Watanabe, Takaaki Hori, John R. Hershey |
| 2017 | ICASSP | Deep clustering and conventional networks for music separation: Stronger together. | Yi Luo, Zhuo Chen, John R. Hershey, Jonathan Le Roux, Nima Mesgarani |
| 2017 | ICASSP | Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognition. | Zhong Meng, Shinji Watanabe, John R. Hershey, Hakan Erdogan |
| 2017 | ICASSP | Student-teacher network learning with enhanced features. | Shinji Watanabe, Takaaki Hori, Jonathan Le Roux, John R. Hershey |
| 2017 | ICCV | Attention-Based Multimodal Fusion for Video Description. | Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, Kazuhiro Sumi |
| 2017 | ICML | Multichannel End-to-end Speech Recognition. | Tsubasa Ochiai, Shinji Watanabe, Takaaki Hori, John R. Hershey |
| 2016 | ICASSP | Deep clustering: Discriminative embeddings for segmentation and separation. | John R. Hershey, Zhuo Chen, Jonathan Le Roux, Shinji Watanabe |
| 2016 | ICASSP | Minimum word error training of long short-term memory recurrent neural network language models for speech recognition. | Takaaki Hori, Chiori Hori, Shinji Watanabe, John R. Hershey |
| 2016 | ICASSP | Deep unfolding for multichannel source separation. | Scott Wisdom, John R. Hershey, Jonathan Le Roux, Shinji Watanabe |
| 2016 | ICASSP | Deep beamforming networks for multi-channel speech recognition. | Xiong Xiao, Shinji Watanabe, Hakan Erdogan, Liang Lu, John R. Hershey, Michael L. Seltzer, Guoguo Chen, Yu Zhang, Michael I. Mandel, Dong Yu |
| 2016 | Interspeech | Improved MVDR Beamforming Using Single-Channel Mask Prediction Networks. | Hakan Erdogan, John R. Hershey, Shinji Watanabe, Michael I. Mandel, Jonathan Le Roux |
| 2016 | Interspeech | Context-Sensitive and Role-Dependent Spoken Language Understanding Using Bidirectional and Attention LSTMs. | Chiori Hori, Takaaki Hori, Shinji Watanabe, John R. Hershey |
| 2016 | Interspeech | Single-Channel Multi-Speaker Separation Using Deep Clustering. | Yusuf Ziya Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, John R. Hershey |
| 2015 | ASRU | The MERL/SRI system for the 3RD CHiME challenge using beamforming, robust feature extraction, and advanced speech recognition. | Takaaki Hori, Zhuo Chen, Hakan Erdogan, John R. Hershey, Jonathan Le Roux, Vikramjit Mitra, Shinji Watanabe |
| 2015 | ICASSP | Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks. | Hakan Erdogan, John R. Hershey, Shinji Watanabe, Jonathan Le Roux |
| 2015 | ICASSP | Deep NMF for speech separation. | Jonathan Le Roux, John R. Hershey, Felix Weninger |
| 2015 | ICASSP | Micbots: Collecting large realistic datasets for speech and audio research using mobile robots. | Jonathan Le Roux, Emmanuel Vincent, John R. Hershey, Daniel P. W. Ellis |
| 2015 | Interspeech | Uncertainty propagation through deep neural networks. | Ahmed Hussen Abdelaziz, Shinji Watanabe, John R. Hershey, Emmanuel Vincent, Dorothea Kolossa |
| 2015 | Interspeech | Speech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networks. | Zhuo Chen, Shinji Watanabe, Hakan Erdogan, John R. Hershey |
| 2014 | ICASSP | Non-negative source-filter dynamical system for speech enhancement. | Umut Simsekli, Jonathan Le Roux, John R. Hershey |
| 2014 | ICASSP | Log-linear dialog manager. | Hao Tang, Shinji Watanabe, Tim K. Marks, John R. Hershey |
| 2014 | Interspeech | Sequential maximum mutual information linear discriminant analysis for speech recognition. | Yuuki Tachioka, Shinji Watanabe, Jonathan Le Roux, John R. Hershey |
| 2014 | Interspeech | Cost-level integration of statistical and rule-based dialog managers. | Shinji Watanabe, John R. Hershey, Tim K. Marks, Youichi Fujii, Yusuke Koji |
| 2014 | Interspeech | Discriminative NMF and its application to single-channel source separation. | Felix Weninger, Jonathan Le Roux, John R. Hershey, Shinji Watanabe |
| 2013 | ASRU | A generalized discriminative training framework for system combination. | Yuuki Tachioka, Shinji Watanabe, Jonathan Le Roux, John R. Hershey |
| 2013 | ICASSP | Non-negative dynamical system with application to speech and audio. | Cdric Fvotte, Jonathan Le Roux, John R. Hershey |
| 2013 | ICASSP | Source localization in reverberant environments using sparse optimization. | Jonathan Le Roux, Petros T. Boufounos, Kang Kang, John R. Hershey |
| 2013 | ICASSP | Effectiveness of discriminative training and feature transformation for reverberated and noisy speech. | Yuuki Tachioka, Shinji Watanabe, John R. Hershey |
| 2013 | ICASSP | Stereo-based feature enhancement using dictionary learning. | Shinji Watanabe, John R. Hershey |
| 2013 | IJCNLP | Statistical Dialogue Management using Intention Dependency Graph. | Koichiro Yoshino, Shinji Watanabe, Jonathan Le Roux, John R. Hershey |
| 2012 | ICASSP | Indirect model-based speech enhancement. | Jonathan Le Roux, John R. Hershey |
| 2011 | ICASSP | Clustering of bootstrapped acoustic model with full covariance. | Xin Chen, Xiaodong Cui, Jian Xue, Peder A. Olsen, John R. Hershey, Bowen Zhou, Yunxin Zhao |
| 2011 | Interspeech | Acoustic Modeling with Bootstrap and Restructuring Based on Full Covariance. | Xiaodong Cui, Xin Chen, Jian Xue, Peder A. Olsen, John R. Hershey, Bowen Zhou |
| 2011 | IROS | Entropy-based motion selection for touch-based registration using Rao-Blackwellized particle filtering. | Yuichi Taguchi, Tim K. Marks, John R. Hershey |
| 2010 | Interspeech | Restructuring exponential family mixture models. | Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen |
| 2010 | Interspeech | Signal interaction and the devil function. | John R. Hershey, Peder A. Olsen, Steven J. Rennie |
| 2010 | Interspeech | Modeling posterior probabilities using the linear exponential family. | Peder A. Olsen, Vaibhava Goel, Charles A. Micchelli, John R. Hershey |
| 2009 | ASRU | Hierarchical variational loopy belief propagation for multi-talker speech recognition. | Steven J. Rennie, John R. Hershey, Peder A. Olsen |
| 2009 | ICASSP | A fast, accurate approximation to log likelihood of Gaussian mixture models. | Pierre L. Dognin, Vaibhava Goel, John R. Hershey, Peder A. Olsen |
| 2009 | ICASSP | Refactoring acoustic models using variational density approximation. | Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen |
| 2009 | ICASSP | Single-channel speech separation and recognition using loopy belief propagation. | Steven J. Rennie, John R. Hershey, Peder A. Olsen |
| 2009 | Interspeech | Refactoring acoustic models using variational expectation-maximization. | Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen |
| 2009 | Interspeech | Variational loopy belief propagation for multi-talker speech recognition. | Steven J. Rennie, John R. Hershey, Peder A. Olsen |
| 2008 | ICASSP | Accelerated Monte Carlo for Kullback-Leibler divergence between Gaussian mixture models. | Jia-Yu Chen, John R. Hershey, Peder A. Olsen, Emmanuel Yashchin |
| 2008 | ICASSP | Variational Bhattacharyya divergence for hidden Markov models. | John R. Hershey, Peder A. Olsen |
| 2008 | ICASSP | Optimizing speech recognition grammars using a measure of similarity between hidden Markov models. | Binit Mohanty, John R. Hershey, Peder A. Olsen, Suleyman Serdar Kozat, Vaibhava Goel |
| 2008 | ICASSP | Efficient model-based speech separation and denoising using non-negative subspace analysis. | Steven J. Rennie, John R. Hershey, Peder A. Olsen |
| 2007 | ASRU | Variational Kullback-Leibler divergence for Hidden Markov models. | John R. Hershey, Peder A. Olsen, Steven J. Rennie |
| 2007 | ICASSP | Approximating the Kullback Leibler Divergence Between Gaussian Mixture Models. | John R. Hershey, Peder A. Olsen |
| 2007 | Interspeech | Word confusability - measuring hidden Markov model similarity. | Jia-Yu Chen, Peder A. Olsen, John R. Hershey |
| 2007 | Interspeech | Bhattacharyya error and divergence using variational importance sampling. | Peder A. Olsen, John R. Hershey |
| 2006 | Interspeech | Super-human multi-talker speech recognition: the IBM 2006 speech separation challenge system. | Trausti T. Kristjansson, John R. Hershey, Peder A. Olsen, Steven J. Rennie, Ramesh A. Gopinath |
| 2006 | Interspeech | The Iroquois model: using temporal dynamics to separate speakers. | Steven J. Rennie, Peder A. Olsen, John R. Hershey, Trausti T. Kristjansson |
| 2004 | CVPR | 3D Tracking of Morphable Objects Using Conditionally Gaussian Nonlinear Filters. | Tim K. Marks, John R. Hershey, J. Cooper Roddey, Javier R. Movellan |
| 2004 | ECCV | Stereo Based 3D Tracking and Scene Learning, Employing Particle Filtering within EM. | Trausti T. Kristjansson, Hagai Attias, John R. Hershey |
| 2004 | ICASSP | Audio-visual graphical models for speech processing. | John R. Hershey, Hagai Attias, Nebojsa Jojic, Trausti T. Kristjansson |
| 2004 | ICASSP | Single microphone source separation using high resolution signal reconstruction. | Trausti T. Kristjansson, Hagai Attias, John R. Hershey |
| 2004 | Interspeech | Model-based fusion of bone and air sensors for speech enhancement and robust speech recognition. | John R. Hershey, Trausti T. Kristjansson, Zhengyou Zhang |
| 2000 | ICIP | A Low-Level Cortical Perception Model with Applications to Image Analysis. | Irina F. Gorodnitsky, John R. Hershey |