Skip to content

George Saon

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

113

Venues

8

Active years

1994–2025

Best venue rank

A*

Where they publish

Papers

113 indexed papers, newest first.

YearVenueTitleAuthors
2025ASRUGranite-speech: open-source speech-aware LLMs with strong English ASR capabilities.George Saon, Avihu Dekel, Alexander Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Xiaodong Cui, Zvi Kons
2025ICASSPKnowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASR.Takashi Fukuda, Gakuto Kurata, George Saon
2025ICASSPLLM based Text Generation for Improved Low-resource Speech Recognition Models.Tohru Nagano, Gakuto Kurata, Samuel Thomas, Hong-Kwang Jeff Kuo, Daniel Bolaos, Hyun Jung, George Saon
2025ICASSPA Non-autoregressive Model for Joint STT and TTS.Vishal Sunder, Brian Kingsbury, George Saon, Samuel Thomas, Slava Shechtman, Hagai Aronowitz, Eric Fosler-Lussier, Luis A. Lastras
2025InterspeechExploring the Limits of Conformer CTC-Encoder for Speech Emotion Recognition using Large Language Models.Edmilson da Silva Morais, Hagai Aronowitz, Aharon Satt, Ron Hoory, Avihu Dekel, Brian Kingsbury, George Saon
2024ICASSPSemi-Autoregressive Streaming ASR with Label Context.Siddhant Arora, George Saon, Shinji Watanabe, Brian Kingsbury
2024ICASSPMultiple Representation Transfer from Large Language Models to End-to-End ASR Systems.Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Masayasu Muraoka, George Saon
2024InterspeechExploring the limits of decoder-only models trained on public speech recognition corpora.Ankit Gupta, George Saon, Brian Kingsbury
2023EMNLPSpeech-enriched Memory for Inference-time Adaptation of ASR Models to Word Dictionaries.Ashish R. Mittal, Sunita Sarawagi, Preethi Jyothi, George Saon, Gakuto Kurata
2023ICASSPDiagonal State Space Augmented Transformers for Speech Recognition.George Saon, Ankit Gupta, Xiaodong Cui
2023ICASSPMulti-Speaker Data Augmentation for Improved end-to-end Automatic Speech Recognition.Samuel Thomas, Hong-Kwang Jeff Kuo, George Saon, Brian Kingsbury
2023InterspeechImproving RNN Transducer Acoustic Models for English Conversational Speech Recognition.Xiaodong Cui, George Saon, Brian Kingsbury
2022ICASSPSpeech Recognition Using Biologically-Inspired Neural Networks.Thomas Bohnstingl, Ayush Garg, Stanislaw Wozniak, George Saon, Evangelos Eleftheriou, Angeliki Pantazi
2022ICASSPImproving End-to-end Models for Set Prediction in Spoken Language Understanding.Hong-Kwang Jeff Kuo, Zoltn Tske, Samuel Thomas, Brian Kingsbury, George Saon
2022ICASSPTowards Reducing the Need for Speech Training Data to Build Spoken Language Understanding Systems.Samuel Thomas, Hong-Kwang Jeff Kuo, Brian Kingsbury, George Saon
2022ICASSPIntegrating Text Inputs for Training and Adapting RNN Transducer ASR Models.Samuel Thomas, Brian Kingsbury, George Saon, Hong-Kwang Jeff Kuo
2022InterspeechImproving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label Smoothing.Xiaodong Cui, George Saon, Tohru Nagano, Masayuki Suzuki, Takashi Fukuda, Brian Kingsbury, Gakuto Kurata
2022InterspeechAccelerating Inference and Language Model Fusion of Recurrent Neural Network Transducers via End-to-End 4-bit Quantization.Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Swagath Venkataramani, George Saon, Xiaodong Cui, Brian Kingsbury, Kailash Gopalakrishnan
2022InterspeechGlobal RNN Transducer Models For Multi-dialect Speech Recognition.Takashi Fukuda, Samuel Thomas, Masayuki Suzuki, Gakuto Kurata, George Saon, Brian Kingsbury
2022InterspeechExtending RNN-T-based speech recognition systems with emotion and language classification.Zvi Kons, Hagai Aronowitz, Edmilson da Silva Morais, Matheus Damasceno, Hong-Kwang Kuo, Samuel Thomas, George Saon
2022InterspeechVQ-T: RNN Transducers using Vector-Quantized Prediction Network States.Jiatong Shi, George Saon, David Haws, Shinji Watanabe, Brian Kingsbury
2022InterspeechEffect and Analysis of Large-scale Language Model Rescoring on Competitive ASR Systems.Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Nobuyasu Itoh, George Saon
2021ICASSPRNN Transducer Models for Spoken Language Understanding.Samuel Thomas, Hong-Kwang Jeff Kuo, George Saon, Zoltn Tske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory
2021ICASSPAdvancing RNN Transducer Technology for Speech Recognition.George Saon, Zoltn Tske, Daniel Bolaos, Brian Kingsbury
2021InterspeechReducing Exposure Bias in Training Recurrent Neural Network Transducers.Xiaodong Cui, Brian Kingsbury, George Saon, David Haws, Zoltn Tske
2021Interspeech4-Bit Quantization of LSTM-Based Speech Recognition Models.Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Xiao Sun, Naigang Wang, Swagath Venkataramani, George Saon, Xiaodong Cui, Brian Kingsbury, Wei Zhang, Zoltn Tske, Kailash Gopalakrishnan
2021InterspeechIntegrating Dialog History into End-to-End Spoken Language Understanding Systems.Jatin Ganhotra, Samuel Thomas, Hong-Kwang Jeff Kuo, Sachindra Joshi, George Saon, Zoltn Tske, Brian Kingsbury
2021InterspeechImproving Customization of Neural Transducers by Mitigating Acoustic Mismatch of Synthesized Audio.Gakuto Kurata, George Saon, Brian Kingsbury, David Haws, Zoltn Tske
2021InterspeechOn the Limit of English Conversational Speech Recognition.Zoltn Tske, George Saon, Brian Kingsbury
2020ICASSPAlignment-Length Synchronous Decoding for RNN Transducer.George Saon, Zoltn Tske, Kartik Audhkhasi
2020ICASSPImproving Efficiency in Large-Scale Decentralized Distributed Training.Wei Zhang, Xiaodong Cui, Abdullah Kayi, Mingrui Liu, Ulrich Finkler, Brian Kingsbury, George Saon, Youssef Mroueh, Alper Buyuktosunoglu, Payel Das, David S. Kung, Michael Picheny
2020InterspeechKnowledge Distillation from Offline to Streaming RNN Transducer for End-to-End Speech Recognition.Gakuto Kurata, George Saon
2020InterspeechSingle Headed Attention Based Sequence-to-Sequence Model for State-of-the-Art Results on Switchboard.Zoltn Tske, George Saon, Kartik Audhkhasi, Brian Kingsbury
2019ASRUSimplified LSTMS for Speech Recognition.George Saon, Zoltn Tske, Kartik Audhkhasi, Brian Kingsbury, Michael Picheny, Samuel Thomas
2019ICASSPSequence Noise Injected Training for End-to-end Speech Recognition.George Saon, Zoltn Tske, Kartik Audhkhasi, Brian Kingsbury
2019ICASSPEnglish Broadcast News Speech Recognition by Humans and Machines.Samuel Thomas, Masayuki Suzuki, Yinghui Huang, Gakuto Kurata, Zoltn Tske, George Saon, Brian Kingsbury, Michael Picheny, Tom Dibert, Alice Kaiser-Schatzlein, Bern Samko
2019ICASSPDistributed Deep Learning Strategies for Automatic Speech Recognition.Wei Zhang, Xiaodong Cui, Ulrich Finkler, Brian Kingsbury, George Saon, David S. Kung, Michael Picheny
2019InterspeechForget a Bit to Learn Better: Soft Forgetting for CTC-Based Automatic Speech Recognition.Kartik Audhkhasi, George Saon, Zoltn Tske, Brian Kingsbury, Michael Picheny
2019InterspeechChallenging the Boundaries of Speech Recognition: The MALACH Corpus.Michael Picheny, Zoltn Tske, Brian Kingsbury, Kartik Audhkhasi, Xiaodong Cui, George Saon
2019InterspeechAdvancing Sequence-to-Sequence Based Speech Recognition.Zoltn Tske, Kartik Audhkhasi, George Saon
2019InterspeechA Highly Efficient Distributed Deep Learning System for Automatic Speech Recognition.Wei Zhang, Xiaodong Cui, Ulrich Finkler, George Saon, Abdullah Kayi, Alper Buyuktosunoglu, Brian Kingsbury, David S. Kung, Michael Picheny
2018ICASSPBuilding Competitive Direct Acoustics-to-Word Models for English Conversational Speech Recognition.Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Michael Picheny
2017ASRULanguage modeling with highway LSTM.Gakuto Kurata, Bhuvana Ramabhadran, George Saon, Abhinav Sethy
2017ICASSPKnowledge distillation across ensembles of multilingual models for low-resource languages.Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Tom Sercu, Kartik Audhkhasi, Abhinav Sethy, Markus Nubaum-Thom, Andrew Rosenberg
2017ICASSPNetwork architectures for multilingual speech representation learning.Tom Sercu, George Saon, Jia Cui, Xiaodong Cui, Bhuvana Ramabhadran, Brian Kingsbury, Abhinav Sethy
2017InterspeechDirect Acoustics-to-Word Models for English Conversational Speech Recognition.Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon, Michael Picheny, David Nahamoo
2017InterspeechEmbedding-Based Speaker Adaptive Training of Deep Neural Networks.Xiaodong Cui, Vaibhava Goel, George Saon
2017InterspeechEmpirical Exploration of Novel Architectures and Objectives for Language Models.Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, George Saon
2017InterspeechEnglish Conversational Telephone Speech Recognition by Humans and Machines.George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas, Dimitrios Dimitriadis, Xiaodong Cui, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall
2017SCAccelerating deep neural network learning for speech recognition on a cluster of GPUs.Guojing Cong, Brian Kingsbury, Soumyadip Gosh, George Saon, Fan Zhou
2016ICASSPOn the importance of event detection for ASR.David Haws, Dimitrios Dimitriadis, George Saon, Samuel Thomas, Michael Picheny
2016InterspeechThe IBM 2016 English Conversational Telephone Speech Recognition System.George Saon, Tom Sercu, Steven J. Rennie, Hong-Kwang Jeff Kuo
2016InterspeechDomain Adaptation of CNN Based Acoustic Models Under Limited Resource Settings.Masayuki Suzuki, Ryuki Tachibana, Samuel Thomas, Bhuvana Ramabhadran, George Saon
2015ICASSPA nonmonotone learning rate strategy for SGD training of deep neural networks.Nitish Shirish Keskar, George Saon
2015ICASSPOrder-free spoken term detection.Lidia Mangu, George Saon, Michael Picheny, Brian Kingsbury
2015ICASSPImprovements to the IBM speech activity detection system for the DARPA RATS program.Samuel Thomas, George Saon, Maarten Van Segbroeck, Shrikanth S. Narayanan
2015InterspeechA multi-region deep neural network model in speech recognition.Jia Cui, George Saon, Bhuvana Ramabhadran, Brian Kingsbury
2015InterspeechThe IBM 2015 English conversational telephone speech recognition system.George Saon, Hong-Kwang Jeff Kuo, Steven J. Rennie, Michael Picheny
2015InterspeechThe IBM BOLT speech transcription system.Samuel Thomas, George Saon, Hong-Kwang Jeff Kuo, Lidia Mangu
2014ICASSPImprovements to filterbank and delta learning within a deep neural network framework.Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George Saon, Bhuvana Ramabhadran
2014ICASSPA comparison of two optimization techniques for sequence discriminative training of deep neural networks.George Saon, Hagen Soltau
2014ICASSPJoint training of convolutional and non-convolutional neural networks.Hagen Soltau, George Saon, Tara N. Sainath
2014ICASSPAnalyzing convolutional neural networks for speech activity detection in mismatched acoustic conditions.Samuel Thomas, Sriram Ganapathy, George Saon, Hagen Soltau
2014InterspeechParallel deep neural network training for LVCSR tasks using blue gene/Q.Tara N. Sainath, I-Hsin Chung, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Brian Kingsbury, George Saon, Vernon Austel, Upendra V. Chaudhari
2014InterspeechUnfolded recurrent neural networks for speech recognition.George Saon, Hagen Soltau, Ahmad Emami, Michael Picheny
2013ASRUThe IBM keyword search system for the DARPA RATS program.Lidia Mangu, Hagen Soltau, Hong-Kwang Kuo, George Saon
2013ASRUImprovements to Deep Convolutional Neural Networks for LVCSR.Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George E. Dahl, George Saon, Hagen Soltau, Toms Beran, Aleksandr Y. Aravkin, Bhuvana Ramabhadran
2013ASRUSpeaker adaptation of neural network acoustic models using i-vectors.George Saon, Hagen Soltau, David Nahamoo, Michael Picheny
2013ICASSPExploiting diversity for spoken term detection.Lidia Mangu, Hagen Soltau, Hong-Kwang Kuo, Brian Kingsbury, George Saon
2013InterspeechThe IBM speech activity detection system for the DARPA RATS program.George Saon, Samuel Thomas, Hagen Soltau, Sriram Ganapathy, Brian Kingsbury
2013InterspeechNeural network acoustic models for the DARPA RATS program.Hagen Soltau, Hong-Kwang Kuo, Lidia Mangu, George Saon, Toms Beran
2012InterspeechSparse Bayesian Factor Analysis for Stereo-based Stochastic Mapping.Xiaodong Cui, Mohamed Afify, George Saon, Vaibhava Goel
2012InterspeechDiscriminative feature-space transforms using deep neural networks.George Saon, Brian Kingsbury
2011ASRUMinimum Bayes risk discriminative language models for Arabic speech recognition.Hong-Kwang Jeff Kuo, Ebru Arisoy, Lidia Mangu, George Saon
2011ASRUThe IBM 2011 GALE Arabic speech transcription system.Lidia Mangu, Hong-Kwang Kuo, Stephen M. Chu, Brian Kingsbury, George Saon, Hagen Soltau, Fadi Biadsy
2011ASRUSome properties of Bayesian sensing hidden Markov models.George Saon, Jen-Tzung Chien
2011ICASSPThe IBM 2009 GALE Arabic speech transcription system.Brian Kingsbury, Hagen Soltau, George Saon, Stephen M. Chu, Hong-Kwang Kuo, Lidia Mangu, Suman V. Ravuri, Nelson Morgan, Adam Janin
2011ICASSPBayesian sensing hidden Markov models for speech recognition.George Saon, Jen-Tzung Chien
2011ICASSPDiscriminative training for Bayesian sensing hidden Markov models.George Saon, Jen-Tzung Chien
2010ICASSPThe IBM 2008 GALE Arabic speech transcription system.George Saon, Hagen Soltau, Upendra V. Chaudhari, Stephen M. Chu, Brian Kingsbury, Hong-Kwang Kuo, Lidia Mangu, Daniel Povey
2010InterspeechBoosting systems for LVCSR.George Saon, Hagen Soltau
2009ASRUDynamic network decoding revisited.Hagen Soltau, George Saon
2009ICASSPLarge margin semi-tied covariance transforms for discriminative training.George Saon, Daniel Povey, Hagen Soltau
2008ICASSPBoosted MMI for model and feature-space discriminative training.Daniel Povey, Dimitri Kanevsky, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Karthik Visweswariah
2008InterspeechPenalty function maximization for large margin HMM training.George Saon, Daniel Povey
2007ASRULattice-based Viterbi decoding techniques for speech translation.George Saon, Michael Picheny
2007ICASSPThe IBM 2006 Gale Arabic ASR System.Hagen Soltau, George Saon, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey, Geoffrey Zweig
2006ICASSPA Non-Linear Speaker Adaptation Technique using Kernel Ridge Regression.George Saon
2006ICASSPAutomated Quality Monitoring in the Call Center with ASR and Maximum Entropy.Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury
2006InterspeechFeature and model space speaker adaptation with full covariance Gaussians.Daniel Povey, George Saon
2006NAACLAutomated Quality Monitoring for Call Centers using Speech and NLP Technologies.Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury
2005ICASSPfMPE: Discriminatively Trained Features for Speech Recognition.Daniel Povey, Brian Kingsbury, Lidia Mangu, George Saon, Hagen Soltau, Geoffrey Zweig
2005ICASSPThe IBM 2004 Conversational Telephony System for Rich Transcription.Hagen Soltau, Brian Kingsbury, Lidia Mangu, Daniel Povey, George Saon, Geoffrey Zweig
2005InterspeechAnatomy of an extremely fast LVCSR decoder.George Saon, Daniel Povey, Geoffrey Zweig
2004ICASSPFeature space Gaussianization.George Saon, Satya Dharanipragada, Daniel Povey
2004ICASSPFractional Fourier transform features for speech recognition.Ruhi Sarikaya, Yuqing Gao, George Saon
2003InterspeechToward domain-independent conversational speech recognition.Brian Kingsbury, Lidia Mangu, George Saon, Geoffrey Zweig, Scott Axelrod, Vaibhava Goel, Karthik Visweswariah, Michael Picheny
2003InterspeechAn architecture for rapid decoding of large vocabulary conversational speech.George Saon, Geoffrey Zweig, Brian Kingsbury, Lidia Mangu, Upendra V. Chaudhari
2002ICASSPDigit recognition in noisy environments via a sequential GMM/SVM system.Shai Fine, George Saon, Ramesh A. Gopinath
2002ICASSPRobust speech recognition in Noisy Environments: The 2001 IBM spine evaluation system.Brian Kingsbury, George Saon, Lidia Mangu, Mukund Padmanabhan, Ruhi Sarikaya
2002InterspeechImprovements to the IBM Aurora 2 multi-condition system.George Saon, Juan M. Huerta
2002InterspeechArc minimization in finite state decoding graphs with cross-word acoustic context.Geoffrey Zweig, George Saon, Franois Yvon
2001ICASSPSpeech recognition for DARPA Communicator.Andrew Aaron, Scott Saobing Chen, Paul S. Cohen, Satya Dharanipragada, Ellen Eide, Martin Franz, Jean-Michel LeRoux, X. Luo, Benot Maison, Lidia Mangu, T. Mathes, Miroslav Novak, Peder A. Olsen, Michael Picheny, Harry Printz, Bhuvana Ramabhadran, Andrej Sakrajda, George Saon, Borivoj Tydlitt, Karthik Visweswariah, D. Yuk
2001ICASSPLinear feature space projections for speaker adaptation.George Saon, Geoffrey Zweig, Mukund Padmanabhan
2001InterspeechRobust digit recognition in noisy environments: the IBM Aurora 2 system.George Saon, Juan M. Huerta, Ea-Ee Jan
2000ICASSPMaximum likelihood discriminant feature spaces.George Saon, Mukund Padmanabhan, Ramesh A. Gopinath, Scott Saobing Chen
2000InterspeechRecent improvements in speech recognition performance on large vocabulary conversational speech (voicemail and switchboard).Jing Huang, Brian Kingsbury, Lidia Mangu, Mukund Padmanabhan, George Saon, Geoffrey Zweig
2000InterspeechReal-time multilingual HMM training robust to channel variations.Ea-Ee Jan, Jaime Botella Ordinas, George Saon, Salim Roukos
2000InterspeechMinimum Bayes error feature selection.George Saon, Mukund Padmanabhan
1999InterspeechRecent improvements in voicemail transcription.Mukund Padmanabhan, George Saon, Sankar Basu, Jing Huang, Geoffrey Zweig
1997ICASSPBinary pattern recognition using Markov random fields and HMMs.George Saon, Abdel Belad
1995ICDARStochastic trajectory modeling for recognition of unconstrained handwritten words.George Saon, Abdel Belad, Yifan Gong
1994MVAOff-line Handwriting Recognition by Statistical Correlation.George Saon, Abdel Belad, Yifan Gong