| 2025 | ICASSP | Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition. | Takaaki Hori, Martin Kocour, Adnan Haider, Erik McDermott, Xiaodan Zhuang |
| 2023 | ICASSP | Neural Transducer Training: Reduced Memory Consumption with Sample-Wise Computation. | Stefan Braun, Erik McDermott, Roger Hsiao |
| 2023 | ICASSP | Variable Attention Masking for Configurable Transformer Transducer Speech Recognition. | Pawel Swietojanski, Stefan Braun, Dogan Can, Thiago Fraga da Silva, Arnab Ghoshal, Takaaki Hori, Roger Hsiao, Henry Mason, Erik McDermott, Honza Silovsky, Ruchir Travadi, Xiaodan Zhuang |
| 2020 | ICASSP | Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss. | Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, Shankar Kumar |
| 2019 | ASRU | A Density Ratio Approach to Language Model Fusion in End-to-End Automatic Speech Recognition. | Erik McDermott, Hasim Sak, Ehsan Variani |
| 2018 | ICASSP | Sampled Connectionist Temporal Classification. | Ehsan Variani, Tom Bagby, Kamel Lahouel, Erik McDermott, Michiel Bacchiani |
| 2017 | Interspeech | Acoustic Modeling for Google Home. | Bo Li, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim, Olivier Siohan, Mitchel Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
| 2017 | Interspeech | End-to-End Training of Acoustic Models for Large Vocabulary Continuous Speech Recognition with TensorFlow. | Ehsan Variani, Tom Bagby, Erik McDermott, Michiel Bacchiani |
| 2015 | ICASSP | A Gaussian Mixture Model layer jointly optimized with discriminative features within a Deep Neural Network architecture. | Ehsan Variani, Erik McDermott, Georg Heigold |
| 2014 | ICASSP | Asynchronous stochastic optimization for sequence training of deep neural networks. | Georg Heigold, Erik McDermott, Vincent Vanhoucke, Andrew W. Senior, Michiel Bacchiani |
| 2014 | ICASSP | Deep neural networks for small footprint text-dependent speaker verification. | Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lpez-Moreno, Javier Gonzalez-Dominguez |
| 2014 | Interspeech | Asynchronous stochastic optimization for sequence training of deep neural networks: towards big data. | Erik McDermott, Georg Heigold, Pedro J. Moreno, Andrew W. Senior, Michiel Bacchiani |
| 2014 | Interspeech | Sequence discriminative distributed training of long short-term memory recurrent neural networks. | Hasim Sak, Oriol Vinyals, Georg Heigold, Andrew W. Senior, Erik McDermott, Rajat Monga, Mark Z. Mao |
| 2013 | ASRU | Large scale deep neural network acoustic modeling with semi-supervised training data for YouTube video transcription. | Hank Liao, Erik McDermott, Andrew W. Senior |
| 2010 | ICASSP | Discriminative training based on an integrated view of MPE and MMI in margin and error space. | Erik McDermott, Shinji Watanabe, Atsushi Nakamura |
| 2010 | ICASSP | A discriminative model for continuous speech recognition based on Weighted Finite State Transducers. | Shinji Watanabe, Takaaki Hori, Erik McDermott, Atsushi Nakamura |
| 2010 | ICASSP | Minimum Error Classification with geometric margin control. | Hideyuki Watanabe, Shigeru Katagiri, Kouta Yamada, Erik McDermott, Atsushi Nakamura, Shinji Watanabe, Miho Ohsaki |
| 2009 | ICASSP | A unified view for discriminative objective functions based on negative exponential of difference measure between strings. | Atsushi Nakamura, Erik McDermott, Shinji Watanabe, Shigeru Katagiri |
| 2009 | Interspeech | Margin-space integration of MPE loss via differencing of MMI functionals for generalized error-weighted discriminative training. | Erik McDermott, Shinji Watanabe, Atsushi Nakamura |
| 2008 | Interspeech | Flexible discriminative training based on equal error group scores obtained from an error-indexed forward-backward algorithm. | Erik McDermott, Atsushi Nakamura |
| 2007 | Interspeech | Discriminative MCE-based speaker adaptation of acoustic models for a spoken lecture processing task. | Timothy J. Hazen, Erik McDermott |
| 2007 | Interspeech | String and lattice based discriminative training for the corpus of spontaneous Japanese lecture transcription task. | Erik McDermott, Atsushi Nakamura |
| 2006 | ACL | Training Conditional Random Fields with Multivariate Evaluation Measures. | Jun Suzuki, Erik McDermott, Hideki Isozaki |
| 2005 | ICASSP | Minimum Classification Error for Large Scale Speech Recognition Tasks using Weighted Finite State Transducers. | Erik McDermott, Shigeru Katagiri |
| 2005 | Interspeech | Optimization methods for discriminative training. | Jonathan Le Roux, Erik McDermott |
| 2004 | ICASSP | Minimum classification error training of landmark models for real-time continuous speech recognition. | Erik McDermott, Timothy J. Hazen |
| 2004 | Interspeech | A theoretical analysis of speech recognition based on feature trajectory models. | Yasuhiro Minami, Erik McDermott, Atsushi Nakamura, Shigeru Katagiri |
| 2003 | ICASSP | A new formalization of minimum classification error using a Parzen estimate of classification chance. | Erik McDermott, Shigeru Katagiri |
| 2003 | ICASSP | Recognition method with parametric trajectory generated from mixture distribution HMMs. | Yasuhiro Minami, Erik McDermott, Atsushi Nakamura, Shigeru Katagiri |
| 2003 | ICASSP | Pervasive unsupervised adaptation for lecture speech transcription. | Daniel Willett, Thomas Niesler, Erik McDermott, Yasuhiro Minami, Shigeru Katagiri |
| 2003 | Interspeech | Blind inversion of multidimensional functions for speech enhancement. | John Hogden, Patrick Valdez, Shigeru Katagiri, Erik McDermott |
| 2002 | ICASSP | A recognition method with parametric trajectory synthesized using direct relations between static and dynamic feature vector time series. | Yasuhiro Minami, Erik McDermott, Atsushi Nakamura, Shigeru Katagiri |
| 2002 | Interspeech | Classification error from the theoretical Bayes classification risk. | Erik McDermott, Shigeru Katagiri |
| 2001 | Interspeech | Time and memory efficient viterbi decoding for LVCSR using a precompiled search network. | Daniel Willett, Erik McDermott, Yasuhiro Minami, Shigeru Katagiri |
| 2000 | ICASSP | Discriminative training for large vocabulary telephone-based name recognition. | Erik McDermott, Alain Biem, Seiichi Tenpaku, Shigeru Katagiri |
| 1998 | Interspeech | Computer-based second language production training by using spectrographic representation and HMM-based speech recognition scores. | Reiko Akahane-Yamada, Erik McDermott, Takahiro Adachi, Hideki Kawahara, John S. Pruitt |
| 1997 | ICASSP | Efficient normalization based upon GPD [generalized probabilistic descent]. | Eric A. Woudenberg, Alain Biem, Erik McDermott, Shigeru Katagiri |
| 1997 | Interspeech | String-level MCE for continuous phoneme recognition. | Erik McDermott, Shigeru Katagiri |
| 1996 | ICASSP | A telephone-based directory assistance system adaptively trained using minimum classification error/generalized probabilistic descent. | Erik McDermott, Eric A. Woudenberg, Shigeru Katagiri |
| 1995 | Interspeech | A discriminative filter bank model for speech recognition. | Alain Biem, Erik McDermott, Shigeru Katagiri |
| 1993 | ICASSP | Prototype-based MCE/GPD training for word spotting and connected word recognition. | Erik McDermott, Shigeru Katagiri |
| 1992 | ICASSP | Prototype-based discriminative training for various speech units. | Erik McDermott, Shigeru Katagiri |
| 1991 | ICASSP | Speaker-independent large vocabulary word recognition using an LVQ/HMM hybrid algorithm. | Hitoshi Iwamida, Shigeru Katagiri, Erik McDermott |
| 1990 | ICASSP | A hybrid speech recognition system using HMMs with an LVQ-trained codebook. | Hitoshi Iwamida, Shigeru Katagiri, Erik McDermott, Yoh'ichi Tohkura |
| 1990 | Interspeech | On the robustness of HMM and ANN speech recognition algorithms. | Yasuhiro Minami, Toshiyuki Hanazawa, Hitoshi Iwamida, Erik McDermott, Kiyohiro Shikano, Shigeru Katagiri, Masaona Kagawa |
| 1989 | ICASSP | A new algorithm for representing acoustic feature dynamics. | Shigeru Katagiri, Erik McDermott, Manami Yokota |
| 1989 | ICASSP | Shift-invariant, multi-category phoneme recognition using Kohonen's LVQ2. | Erik McDermott, Shigeru Katagiri |