Skip to content

John R. Hershey

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

94

Venues

12

Active years

2000–2025

Best venue rank

A*

Where they publish

Papers

94 indexed papers, newest first.

YearVenueTitleAuthors
2025ICASSPGenerative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement.Jae-Sung Bae, Anastasia Kuznetsova, Dinesh Manocha, John R. Hershey, Trausti T. Kristjansson, Minje Kim
2025ICASSPTowards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables.Artem Dementyev, Chandan K. A. Reddy, Scott Wisdom, Navin Chatlani, John R. Hershey, Richard F. Lyon
2025ICASSPGenerative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation.Jackie Lin, Georg Gtz, Hermes Sampedro Llopis, Haukur Hafsteinsson, Steinar Gujnsson, Daniel Gert Nielsen, Finnur Pind, Paris Smaragdis, Dinesh Manocha, John R. Hershey, Trausti T. Kristjansson, Minje Kim
2025ICLRI-Con: A Unifying Framework for Representation Learning.Shaden Naif Alshammari, John R. Hershey, Axel Feldmann, William T. Freeman, Mark Hamilton
2024CVPRSeparating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language.Mark Hamilton, Andrew Zisserman, John R. Hershey, William T. Freeman
2024ICASSPUnsupervised Multi-Channel Separation And Adaptation.Cong Han, Kevin W. Wilson, Scott Wisdom, John R. Hershey
2024InterspeechUnsupervised Improved MVDR Beamforming for Sound Enhancement.Jacob Kealey, John R. Hershey, Franois Grondin
2023ICASSPAudioslots: A Slot-Centric Generative Model For Audio Separation.Pradyumna Reddy, Scott Wisdom, Klaus Greff, John R. Hershey, Thomas Kipf
2023InterspeechTokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition.Hakan Erdogan, Scott Wisdom, Xuankai Chang, Zaln Borsos, Marco Tagliasacchi, Neil Zeghidour, John R. Hershey
2022ECCVAudioScopeV2: Audio-Visual Attention Architectures for Calibrated Open-Domain On-Screen Sound Separation.Efthymios Tzinis, Scott Wisdom, Tal Remez, John R. Hershey
2022ICASSPImproving Bird Classification with Unsupervised Sound Separation.Tom Denton, Scott Wisdom, John R. Hershey
2022ICASSPAdapting Speech Separation to Real-World Meetings using Mixture Invariant Training.Aswin Sivaraman, Scott Wisdom, Hakan Erdogan, John R. Hershey
2022InterspeechCycleGAN-based Unpaired Speech Dereverberation.Hannah Muckenhirn, Aleksandr Safin, Hakan Erdogan, Felix de Chaumont Quitry, Marco Tagliasacchi, Scott Wisdom, John R. Hershey
2022InterspeechDistance-Based Sound Separation.Katharine Patterson, Kevin W. Wilson, Scott Wisdom, John R. Hershey
2021ICASSPEnd-To-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings.Soumi Maiti, Hakan Erdogan, Kevin W. Wilson, Scott Wisdom, Shinji Watanabe, John R. Hershey
2021ICASSPSound Event Detection and Separation: A Benchmark on Desed Synthetic Soundscapes.Nicolas Turpault, Romain Serizel, Scott Wisdom, Hakan Erdogan, John R. Hershey, Eduardo Fonseca, Prem Seetharaman, Justin Salamon
2021ICASSPWhat's all the Fuss about Free Universal Sound Separation Data?Scott Wisdom, Hakan Erdogan, Daniel P. W. Ellis, Romain Serizel, Nicolas Turpault, Eduardo Fonseca, Justin Salamon, Prem Seetharaman, John R. Hershey
2021ICLRInto the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds.Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey, Tal Remez, Dan Ellis, John R. Hershey
2021InterspeechContinuous Speech Separation Using Speaker Inventory for Long Recording.Cong Han, Yi Luo, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, Zhuo Chen
2020ICASSPImproving Universal Sound Separation Using Sound Classification.Efthymios Tzinis, Scott Wisdom, John R. Hershey, Aren Jansen, Daniel P. W. Ellis
2019ICASSPSDR - Half-baked or Well Done?Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, John R. Hershey
2019ICASSPThe Phasebook: Building Complex Masks via Discrete Representations for Source Separation.Jonathan Le Roux, Gordon Wichern, Shinji Watanabe, Andy M. Sarroff, John R. Hershey
2019ICASSPDifferentiable Consistency Constraints for Improved Deep Speech Enhancement.Scott Wisdom, John R. Hershey, Kevin W. Wilson, Jeremy Thorpe, Michael Chinen, Brian Patton, Rif A. Saurous
2019InterspeechEnd-to-End Multilingual Multi-Speaker Speech Recognition.Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2019InterspeechVoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking.Quan Wang, Hannah Muckenhirn, Kevin W. Wilson, Prashant Sridhar, Zelin Wu, John R. Hershey, Rif A. Saurous, Ron J. Weiss, Ye Jia, Ignacio Lpez-Moreno
2018ACLA Purely End-to-End System for Multi-speaker Speech Recognition.Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2018ICASSPSpeaker Adaptation for Multichannel End-to-End Speech Recognition.Tsubasa Ochiai, Shinji Watanabe, Shigeru Katagiri, Takaaki Hori, John R. Hershey
2018ICASSPAn End-to-End Language-Tracking Speech Recognizer for Mixed-Language Speech.Hiroshi Seki, Shinji Watanabe, Takaaki Hori, Jonathan Le Roux, John R. Hershey
2018ICASSPEnd-to-End Multi-Speaker Speech Recognition.Shane Settle, Jonathan Le Roux, Takaaki Hori, Shinji Watanabe, John R. Hershey
2018ICASSPMulti-Channel Deep Clustering: Discriminative Spectral and Spatial Embeddings for Speaker-Independent Speech Separation.Zhong-Qiu Wang, Jonathan Le Roux, John R. Hershey
2018ICASSPAlternative Objective Functions for Deep Clustering.Zhong-Qiu Wang, Jonathan Le Roux, John R. Hershey
2018InterspeechEnd-to-End Speech Separation with Unfolded Iterative Phase Reconstruction.Zhong-Qiu Wang, Jonathan Le Roux, DeLiang Wang, John R. Hershey
2017ACLJoint CTC/attention decoding for end-to-end speech recognition.Takaaki Hori, Shinji Watanabe, John R. Hershey
2017ASRUEarly and late integration of audio features for automatic video description.Chiori Hori, Takaaki Hori, Tim K. Marks, John R. Hershey
2017ASRUMulti-level language modeling and decoding for open vocabulary end-to-end speech recognition.Takaaki Hori, Shinji Watanabe, John R. Hershey
2017ASRULanguage independent end-to-end architecture for joint language identification and speech recognition.Shinji Watanabe, Takaaki Hori, John R. Hershey
2017ICASSPDeep clustering and conventional networks for music separation: Stronger together.Yi Luo, Zhuo Chen, John R. Hershey, Jonathan Le Roux, Nima Mesgarani
2017ICASSPDeep long short-term memory adaptive beamforming networks for multichannel robust speech recognition.Zhong Meng, Shinji Watanabe, John R. Hershey, Hakan Erdogan
2017ICASSPStudent-teacher network learning with enhanced features.Shinji Watanabe, Takaaki Hori, Jonathan Le Roux, John R. Hershey
2017ICCVAttention-Based Multimodal Fusion for Video Description.Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, Kazuhiro Sumi
2017ICMLMultichannel End-to-end Speech Recognition.Tsubasa Ochiai, Shinji Watanabe, Takaaki Hori, John R. Hershey
2016ICASSPDeep clustering: Discriminative embeddings for segmentation and separation.John R. Hershey, Zhuo Chen, Jonathan Le Roux, Shinji Watanabe
2016ICASSPMinimum word error training of long short-term memory recurrent neural network language models for speech recognition.Takaaki Hori, Chiori Hori, Shinji Watanabe, John R. Hershey
2016ICASSPDeep unfolding for multichannel source separation.Scott Wisdom, John R. Hershey, Jonathan Le Roux, Shinji Watanabe
2016ICASSPDeep beamforming networks for multi-channel speech recognition.Xiong Xiao, Shinji Watanabe, Hakan Erdogan, Liang Lu, John R. Hershey, Michael L. Seltzer, Guoguo Chen, Yu Zhang, Michael I. Mandel, Dong Yu
2016InterspeechImproved MVDR Beamforming Using Single-Channel Mask Prediction Networks.Hakan Erdogan, John R. Hershey, Shinji Watanabe, Michael I. Mandel, Jonathan Le Roux
2016InterspeechContext-Sensitive and Role-Dependent Spoken Language Understanding Using Bidirectional and Attention LSTMs.Chiori Hori, Takaaki Hori, Shinji Watanabe, John R. Hershey
2016InterspeechSingle-Channel Multi-Speaker Separation Using Deep Clustering.Yusuf Ziya Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, John R. Hershey
2015ASRUThe MERL/SRI system for the 3RD CHiME challenge using beamforming, robust feature extraction, and advanced speech recognition.Takaaki Hori, Zhuo Chen, Hakan Erdogan, John R. Hershey, Jonathan Le Roux, Vikramjit Mitra, Shinji Watanabe
2015ICASSPPhase-sensitive and recognition-boosted speech separation using deep recurrent neural networks.Hakan Erdogan, John R. Hershey, Shinji Watanabe, Jonathan Le Roux
2015ICASSPDeep NMF for speech separation.Jonathan Le Roux, John R. Hershey, Felix Weninger
2015ICASSPMicbots: Collecting large realistic datasets for speech and audio research using mobile robots.Jonathan Le Roux, Emmanuel Vincent, John R. Hershey, Daniel P. W. Ellis
2015InterspeechUncertainty propagation through deep neural networks.Ahmed Hussen Abdelaziz, Shinji Watanabe, John R. Hershey, Emmanuel Vincent, Dorothea Kolossa
2015InterspeechSpeech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networks.Zhuo Chen, Shinji Watanabe, Hakan Erdogan, John R. Hershey
2014ICASSPNon-negative source-filter dynamical system for speech enhancement.Umut Simsekli, Jonathan Le Roux, John R. Hershey
2014ICASSPLog-linear dialog manager.Hao Tang, Shinji Watanabe, Tim K. Marks, John R. Hershey
2014InterspeechSequential maximum mutual information linear discriminant analysis for speech recognition.Yuuki Tachioka, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2014InterspeechCost-level integration of statistical and rule-based dialog managers.Shinji Watanabe, John R. Hershey, Tim K. Marks, Youichi Fujii, Yusuke Koji
2014InterspeechDiscriminative NMF and its application to single-channel source separation.Felix Weninger, Jonathan Le Roux, John R. Hershey, Shinji Watanabe
2013ASRUA generalized discriminative training framework for system combination.Yuuki Tachioka, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2013ICASSPNon-negative dynamical system with application to speech and audio.Cdric Fvotte, Jonathan Le Roux, John R. Hershey
2013ICASSPSource localization in reverberant environments using sparse optimization.Jonathan Le Roux, Petros T. Boufounos, Kang Kang, John R. Hershey
2013ICASSPEffectiveness of discriminative training and feature transformation for reverberated and noisy speech.Yuuki Tachioka, Shinji Watanabe, John R. Hershey
2013ICASSPStereo-based feature enhancement using dictionary learning.Shinji Watanabe, John R. Hershey
2013IJCNLPStatistical Dialogue Management using Intention Dependency Graph.Koichiro Yoshino, Shinji Watanabe, Jonathan Le Roux, John R. Hershey
2012ICASSPIndirect model-based speech enhancement.Jonathan Le Roux, John R. Hershey
2011ICASSPClustering of bootstrapped acoustic model with full covariance.Xin Chen, Xiaodong Cui, Jian Xue, Peder A. Olsen, John R. Hershey, Bowen Zhou, Yunxin Zhao
2011InterspeechAcoustic Modeling with Bootstrap and Restructuring Based on Full Covariance.Xiaodong Cui, Xin Chen, Jian Xue, Peder A. Olsen, John R. Hershey, Bowen Zhou
2011IROSEntropy-based motion selection for touch-based registration using Rao-Blackwellized particle filtering.Yuichi Taguchi, Tim K. Marks, John R. Hershey
2010InterspeechRestructuring exponential family mixture models.Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen
2010InterspeechSignal interaction and the devil function.John R. Hershey, Peder A. Olsen, Steven J. Rennie
2010InterspeechModeling posterior probabilities using the linear exponential family.Peder A. Olsen, Vaibhava Goel, Charles A. Micchelli, John R. Hershey
2009ASRUHierarchical variational loopy belief propagation for multi-talker speech recognition.Steven J. Rennie, John R. Hershey, Peder A. Olsen
2009ICASSPA fast, accurate approximation to log likelihood of Gaussian mixture models.Pierre L. Dognin, Vaibhava Goel, John R. Hershey, Peder A. Olsen
2009ICASSPRefactoring acoustic models using variational density approximation.Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen
2009ICASSPSingle-channel speech separation and recognition using loopy belief propagation.Steven J. Rennie, John R. Hershey, Peder A. Olsen
2009InterspeechRefactoring acoustic models using variational expectation-maximization.Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen
2009InterspeechVariational loopy belief propagation for multi-talker speech recognition.Steven J. Rennie, John R. Hershey, Peder A. Olsen
2008ICASSPAccelerated Monte Carlo for Kullback-Leibler divergence between Gaussian mixture models.Jia-Yu Chen, John R. Hershey, Peder A. Olsen, Emmanuel Yashchin
2008ICASSPVariational Bhattacharyya divergence for hidden Markov models.John R. Hershey, Peder A. Olsen
2008ICASSPOptimizing speech recognition grammars using a measure of similarity between hidden Markov models.Binit Mohanty, John R. Hershey, Peder A. Olsen, Suleyman Serdar Kozat, Vaibhava Goel
2008ICASSPEfficient model-based speech separation and denoising using non-negative subspace analysis.Steven J. Rennie, John R. Hershey, Peder A. Olsen
2007ASRUVariational Kullback-Leibler divergence for Hidden Markov models.John R. Hershey, Peder A. Olsen, Steven J. Rennie
2007ICASSPApproximating the Kullback Leibler Divergence Between Gaussian Mixture Models.John R. Hershey, Peder A. Olsen
2007InterspeechWord confusability - measuring hidden Markov model similarity.Jia-Yu Chen, Peder A. Olsen, John R. Hershey
2007InterspeechBhattacharyya error and divergence using variational importance sampling.Peder A. Olsen, John R. Hershey
2006InterspeechSuper-human multi-talker speech recognition: the IBM 2006 speech separation challenge system.Trausti T. Kristjansson, John R. Hershey, Peder A. Olsen, Steven J. Rennie, Ramesh A. Gopinath
2006InterspeechThe Iroquois model: using temporal dynamics to separate speakers.Steven J. Rennie, Peder A. Olsen, John R. Hershey, Trausti T. Kristjansson
2004CVPR3D Tracking of Morphable Objects Using Conditionally Gaussian Nonlinear Filters.Tim K. Marks, John R. Hershey, J. Cooper Roddey, Javier R. Movellan
2004ECCVStereo Based 3D Tracking and Scene Learning, Employing Particle Filtering within EM.Trausti T. Kristjansson, Hagai Attias, John R. Hershey
2004ICASSPAudio-visual graphical models for speech processing.John R. Hershey, Hagai Attias, Nebojsa Jojic, Trausti T. Kristjansson
2004ICASSPSingle microphone source separation using high resolution signal reconstruction.Trausti T. Kristjansson, Hagai Attias, John R. Hershey
2004InterspeechModel-based fusion of bone and air sensors for speech enhancement and robust speech recognition.John R. Hershey, Trausti T. Kristjansson, Zhengyou Zhang
2000ICIPA Low-Level Cortical Perception Model with Applications to Image Analysis.Irina F. Gorodnitsky, John R. Hershey