| 2024 | ICASSP | End-to-End Speech Recognition Contextualization with Large Language Models. | Egor Lakomkin, Chunyang Wu, Yassir Fathullah, Ozlem Kalinli, Michael L. Seltzer, Christian Fuegen |
| 2024 | Interspeech | Towards measuring fairness in speech recognition: Fair-Speech dataset. | Irina-Elena Veliche, Zhuangqun Huang, Vineeth Ayyat Kochaniyan, Fuchun Peng, Ozlem Kalinli, Michael L. Seltzer |
| 2023 | ICASSP | Factorized Blank Thresholding for Improved Runtime Efficiency of Neural Transducers. | Duc Le, Frank Seide, Yuhao Wang, Yang Li, Kjell Schubert, Ozlem Kalinli, Michael L. Seltzer |
| 2023 | ICASSP | Improving fast-slow Encoder based Transducer with Streaming Deliberation. | Ke Li, Jay Mahadeokar, Jinxi Guo, Yangyang Shi, Gil Keren, Ozlem Kalinli, Michael L. Seltzer, Duc Le |
| 2023 | ICASSP | Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities. | Andros Tjandra, Nayan Singhal, David Zhang, Ozlem Kalinli, Abdelrahman Mohamed, Duc Le, Michael L. Seltzer |
| 2023 | Interspeech | Modality Confidence Aware Training for Robust End-to-End Spoken Language Understanding. | Suyoun Kim, Akshat Shrivastava, Duc Le, Ju Lin, Ozlem Kalinli, Michael L. Seltzer |
| 2022 | ICASSP | Neural-FST Class Language Model for End-to-End Speech Recognition. | Antoine Bruguier, Duc Le, Rohit Prabhavalkar, Dangna Li, Zhe Liu, Bo Wang, Eun Chang, Fuchun Peng, Ozlem Kalinli, Michael L. Seltzer |
| 2022 | Interspeech | Evaluating User Perception of Speech Recognition System Quality with Semantic Distance Metric. | Suyoun Kim, Duc Le, Weiyi Zheng, Tarun Singh, Abhinav Arora, Xiaoyu Zhai, Christian Fuegen, Ozlem Kalinli, Michael L. Seltzer |
| 2022 | Interspeech | Deliberation Model for On-Device Spoken Language Understanding. | Duc Le, Akshat Shrivastava, Paden D. Tomasello, Suyoun Kim, Aleksandr Livshits, Ozlem Kalinli, Michael L. Seltzer |
| 2022 | Interspeech | Streaming parallel transducer beam search with fast slow cascaded encoders. | Jay Mahadeokar, Yangyang Shi, Ke Li, Duc Le, Jiedan Zhu, Vikas Chandra, Ozlem Kalinli, Michael L. Seltzer |
| 2021 | ICASSP | Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer. | Suyoun Kim, Yuan Shangguan, Jay Mahadeokar, Antoine Bruguier, Christian Fuegen, Michael L. Seltzer, Duc Le |
| 2021 | ICASSP | Memory-Efficient Speech Recognition on Smart Devices. | Ganesh Venkatesh, Alagappan Valliappan, Jay Mahadeokar, Yuan Shangguan, Christian Fuegen, Michael L. Seltzer, Vikas Chandra |
| 2021 | Interspeech | Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding. | Suyoun Kim, Abhinav Arora, Duc Le, Ching-Feng Yeh, Christian Fuegen, Ozlem Kalinli, Michael L. Seltzer |
| 2021 | Interspeech | Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion. | Duc Le, Mahaveer Jain, Gil Keren, Suyoun Kim, Yangyang Shi, Jay Mahadeokar, Julian Chan, Yuan Shangguan, Christian Fuegen, Ozlem Kalinli, Yatharth Saraf, Michael L. Seltzer |
| 2021 | Interspeech | Flexi-Transducer: Optimizing Latency, Accuracy and Compute for Multi-Domain On-Device Scenarios. | Jay Mahadeokar, Yangyang Shi, Yuan Shangguan, Chunyang Wu, Alex Xiao, Hang Su, Duc Le, Ozlem Kalinli, Christian Fuegen, Michael L. Seltzer |
| 2021 | Interspeech | Collaborative Training of Acoustic Encoders for Speech Recognition. | Varun Nagaraja, Yangyang Shi, Ganesh Venkatesh, Ozlem Kalinli, Michael L. Seltzer, Vikas Chandra |
| 2021 | Interspeech | Dissecting User-Perceived Latency of On-Device E2E Speech Recognition. | Yuan Shangguan, Rohit Prabhavalkar, Hang Su, Jay Mahadeokar, Yangyang Shi, Jiatong Zhou, Chunyang Wu, Duc Le, Ozlem Kalinli, Christian Fuegen, Michael L. Seltzer |
| 2021 | Interspeech | Dynamic Encoder Transducer: A Flexible Solution for Trading Off Accuracy for Latency. | Yangyang Shi, Varun Nagaraja, Chunyang Wu, Jay Mahadeokar, Duc Le, Rohit Prabhavalkar, Alex Xiao, Ching-Feng Yeh, Julian Chan, Christian Fuegen, Ozlem Kalinli, Michael L. Seltzer |
| 2020 | ICASSP | Aipnet: Generative Adversarial Pre-Training of Accent-Invariant Networks for End-To-End Speech Recognition. | Yi-Chen Chen, Zhaojun Yang, Ching-Feng Yeh, Mahaveer Jain, Michael L. Seltzer |
| 2020 | ICASSP | G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR. | Duc Le, Thilo Khler, Christian Fuegen, Michael L. Seltzer |
| 2020 | ICASSP | Transformer-Based Acoustic Modeling for Hybrid Speech Recognition. | Yongqiang Wang, Abdelrahman Mohamed, Duc Le, Chunxi Liu, Alex Xiao, Jay Mahadeokar, Hongzhao Huang, Andros Tjandra, Xiaohui Zhang, Frank Zhang, Christian Fuegen, Geoffrey Zweig, Michael L. Seltzer |
| 2020 | Interspeech | Weak-Attention Suppression for Transformer Based Speech Recognition. | Yangyang Shi, Yongqiang Wang, Chunyang Wu, Christian Fuegen, Frank Zhang, Duc Le, Ching-Feng Yeh, Michael L. Seltzer |
| 2019 | ASRU | From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition. | Duc Le, Xiaohui Zhang, Weiyi Zheng, Christian Fgen, Geoffrey Zweig, Michael L. Seltzer |
| 2019 | ICASSP | End-to-end Contextual Speech Recognition Using Class Language Models and a Token Passing Decoder. | Zhehuai Chen, Mahaveer Jain, Yongqiang Wang, Michael L. Seltzer, Christian Fuegen |
| 2019 | Interspeech | Joint Grapheme and Phoneme Embeddings for Contextual End-to-End ASR. | Zhehuai Chen, Mahaveer Jain, Yongqiang Wang, Michael L. Seltzer, Christian Fuegen |
| 2018 | ICASSP | Efficient Integration of Fixed Beamformers and Speech Separation Networks for Multi-Channel Far-Field Speech Separation. | Zhuo Chen, Takuya Yoshioka, Xiong Xiao, Linyu Li, Michael L. Seltzer, Yifan Gong |
| 2018 | ICASSP | Towards Language-Universal End-to-End Speech Recognition. | Suyoun Kim, Michael L. Seltzer |
| 2018 | Interspeech | Improved Training for Online End-to-end Speech Recognition Systems. | Suyoun Kim, Michael L. Seltzer, Jinyu Li, Rui Zhao |
| 2017 | EACL | May I take your order? A Neural Model for Extracting Structured Information from Conversations. | Baolin Peng, Michael L. Seltzer, Y. C. Ju, Geoffrey Zweig, Kam-Fai Wong |
| 2017 | ICASSP | A study on data augmentation of reverberant speech for robust speech recognition. | Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L. Seltzer, Sanjeev Khudanpur |
| 2017 | Interspeech | Large-Scale Domain Adaptation via Teacher-Student Learning. | Jinyu Li, Michael L. Seltzer, Xi Wang, Rui Zhao, Yifan Gong |
| 2016 | ICASSP | Linearly augmented deep neural network. | Pegah Ghahremani, Jasha Droppo, Michael L. Seltzer |
| 2016 | ICASSP | Deep beamforming networks for multi-channel speech recognition. | Xiong Xiao, Shinji Watanabe, Hakan Erdogan, Liang Lu, John R. Hershey, Michael L. Seltzer, Guoguo Chen, Yu Zhang, Michael I. Mandel, Dong Yu |
| 2016 | Interspeech | On the Role of Nonlinear Transformations in Deep Neural Network Acoustic Models. | Tasha Nagamine, Michael L. Seltzer, Nima Mesgarani |
| 2015 | ICASSP | Improving speech recognition in reverberation using a room-aware deep neural network and multi-task learning. | Ritwik Giri, Michael L. Seltzer, Jasha Droppo, Dong Yu |
| 2015 | ICASSP | Speech recognition with prediction-adaptation-correction recurrent neural networks. | Yu Zhang, Dong Yu, Michael L. Seltzer, Jasha Droppo |
| 2015 | Interspeech | Exploring how deep neural networks form phonemic categories. | Tasha Nagamine, Michael L. Seltzer, Nima Mesgarani |
| 2014 | ICASSP | Factored adaptation of speaker and environment using orthogonal subspace transforms. | Hyunson Seo, Hong-Goo Kang, Michael L. Seltzer |
| 2014 | ICASSP | Single-channel mixed speech recognition using deep neural networks. | Chao Weng, Dong Yu, Michael L. Seltzer, Jasha Droppo |
| 2014 | Interspeech | Towards better performance with heterogeneous training data in acoustic modeling using deep neural networks. | Yan Huang, Malcolm Slaney, Michael L. Seltzer, Yifan Gong |
| 2014 | Interspeech | The influence of pitch and noise on the discriminability of filterbank features. | Malcolm Slaney, Michael L. Seltzer |
| 2014 | Interspeech | An introduction to computational networks and the computational network toolkit (invited talk). | Dong Yu, Adam Eversole, Michael L. Seltzer, Kaisheng Yao, Brian Guenter, Oleksii Kuchaiev, Frank Seide, Huaming Wang, Jasha Droppo, Zhiheng Huang, Geoffrey Zweig, Christopher J. Rossbach, Jon Currey |
| 2013 | ICASSP | Recent advances in deep learning for speech research at Microsoft. | Li Deng, Jinyu Li, Jui-Ting Huang, Kaisheng Yao, Dong Yu, Frank Seide, Michael L. Seltzer, Geoffrey Zweig, Xiaodong He, Jason D. Williams, Yifan Gong, Alex Acero |
| 2013 | ICASSP | Multi-task learning in deep neural networks for improved phoneme recognition. | Michael L. Seltzer, Jasha Droppo |
| 2013 | ICASSP | An investigation of deep neural networks for noise robust speech recognition. | Michael L. Seltzer, Dong Yu, Yongqiang Wang |
| 2013 | ICASSP | Deep neural network features and semi-supervised training for low resource speech recognition. | Samuel Thomas, Michael L. Seltzer, Kenneth Church, Hynek Hermansky |
| 2012 | ICASSP | Improvements to VTS feature enhancement. | Jinyu Li, Michael L. Seltzer, Yifan Gong |
| 2012 | Interspeech | Efficient VTS Adaptation Using Jacobian Approximation. | Jinyu Li, Michael L. Seltzer, Yifan Gong |
| 2012 | Interspeech | Factored adaptation using a combination of feature-space and model-space transforms. | Michael L. Seltzer, Alex Acero |
| 2011 | ASRU | Factored adaptation for separable compensation of speaker and environmental variability. | Michael L. Seltzer, Alex Acero |
| 2011 | ICASSP | Joint encoding of the waveform and speech recognition features using a transform codec. | Xing Fan, Michael L. Seltzer, Jasha Droppo, Henrique S. Malvar, Alex Acero |
| 2011 | ICASSP | CROWDMOS: An approach for crowdsourcing mean opinion score studies. | Flavio P. Ribeiro, Dinei A. F. Florncio, Cha Zhang, Michael L. Seltzer |
| 2011 | Interspeech | Separating Speaker and Environmental Variability Using Factored Transforms. | Michael L. Seltzer, Alex Acero |
| 2011 | Interspeech | Improved Bottleneck Features Using Pretrained Deep Neural Networks. | Dong Yu, Michael L. Seltzer |
| 2010 | ICASSP | Acoustic model adaptation via Linear Spline Interpolation for robust speech recognition. | Michael L. Seltzer, Alex Acero, Kaustubh Kalgaonkar |
| 2010 | Interspeech | Binary coding of speech spectrograms using a deep auto-encoder. | Li Deng, Michael L. Seltzer, Dong Yu, Alex Acero, Abdel-rahman Mohamed, Geoffrey E. Hinton |
| 2010 | Interspeech | HMM adaptation using linear spline interpolation with integrated spline parameter training for robust speech recognition. | Michael L. Seltzer, Alex Acero |
| 2009 | ASRU | Noise robust model adaptation using linear spline interpolation. | Kaustubh Kalgaonkar, Michael L. Seltzer, Alex Acero |
| 2009 | ICASSP | Noise adaptive training using a vector taylor series approach for noise robust automatic speech recognition. | Ozlem Kalinli, Michael L. Seltzer, Alex Acero |
| 2009 | ICASSP | The data deluge: Challenges and opportunities of unlimited data in statistical signal processing. | Michael L. Seltzer, Lei Zhang |
| 2009 | ICASSP | Voice search of structured media data. | Young-In Song, Ye-Yi Wang, Yun-Cheng Ju, Michael L. Seltzer, Ivan Tashev, Alex Acero |
| 2009 | Interspeech | Improving perceived accuracy for in-car media search. | Yun-Cheng Ju, Michael L. Seltzer, Ivan Tashev |
| 2008 | ICASSP | Robust design of wideband loudspeaker arrays. | Ivan Tashev, Jasha Droppo, Michael L. Seltzer, Alex Acero |
| 2008 | ICASSP | Maximum a posteriori ICA: Applying prior knowledge to the separation of acoustic sources. | Graham W. Taylor, Michael L. Seltzer, Alex Acero |
| 2008 | Interspeech | Towards a non-parametric acoustic model: an acoustic decision tree for observation probability calculation. | Jasha Droppo, Michael L. Seltzer, Alex Acero, Yu-Hsiang Bosco Chiu |
| 2007 | ICASSP | Microphone Array Post-Filter using Incremental Bayes Learning to Track the Spatial Distributions of Speech and Noise. | Michael L. Seltzer, Ivan Tashev, Alex Acero |
| 2007 | Interspeech | Robust location understanding in spoken dialog systems using intersections. | Michael L. Seltzer, Yun-Cheng Ju, Ivan Tashev, Alex Acero |
| 2007 | SIGdial | Commute UX: Telephone Dialog System for Location-based Services. | Ivan Tashev, Michael L. Seltzer, Yun-Cheng Ju, Dong Yu, Alex Acero |
| 2006 | Interspeech | Automatic removal of typed keystrokes from speech signals. | Amarnag Subramanya, Michael L. Seltzer, Alex Acero |
| 2005 | ICASSP | Training Wideband Acoustic Models using Mixed-Bandwidth Training Data via Feature Bandwidth Extension. | Michael L. Seltzer, Alex Acero |
| 2005 | Interspeech | Robust bandwidth extension of noise-corrupted narrowband speech. | Michael L. Seltzer, Alex Acero, Jasha Droppo |
| 2004 | ICASSP | Parameter sharing in subband likelihood-maximizing beamforming for speech recognition using microphone arrays. | Michael L. Seltzer, Richard M. Stern |
| 2003 | ICASSP | Subband parameter optimization of microphone arrays for speech recognition in reverberant environments. | Michael L. Seltzer, Richard M. Stern |
| 2003 | Interspeech | A harmonic-model-based front end for robust speech recognition. | Michael L. Seltzer, Jasha Droppo, Alex Acero |
| 2002 | ICASSP | Speech recognizer-based microphone array processing for robust hands-free speech recognition. | Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
| 2001 | ICASSP | Speech in Noisy Environments: robust automatic segmentation, feature extraction, and hypothesis combination. | Rita Singh, Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
| 2001 | Interspeech | Calibration of microphone arrays for improved speech recognition. | Michael L. Seltzer, Bhiksha Raj |
| 2000 | Interspeech | Reconstruction of damaged spectrographic features for robust speech recognition. | Bhiksha Raj, Michael L. Seltzer, Richard M. Stern |
| 2000 | Interspeech | Classifier-based mask estimation for missing feature methods of robust speech recognition. | Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |