Brian Kingsbury
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
131
Venues
9
Active years
1997–2025
Best venue rank
A*
Where they publish
Papers
131 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | ASRU | MADASR 2.0: Multi-Lingual Multi-Dialect ASR Challenge in 8 Indian Languages. | Saurabh Kumar, Sumit Sharma, Deekshitha G, Abhayjeet Singh, Amartyaveer, Sathvik Udupa, Sandhya Badiger, Sanjeev Khudanpur, Sunayana Sitaram, Srinivasan Umesh, Bhuvana Ramabhadran, Brian Kingsbury, Hema A. Murthy, Srikanth S. Narayanan, Howard Lakougna, Prasanta Kumar Ghosh |
| 2025 | ASRU | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities. | George Saon, Avihu Dekel, Alexander Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Xiaodong Cui, Zvi Kons |
| 2025 | CVPR | CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment. | Edson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati, Samuel Thomas, Brian Kingsbury, Leonid Karlinsky, Rogrio Feris, James R. Glass, Hilde Kuehne |
| 2025 | ICASSP | A Non-autoregressive Model for Joint STT and TTS. | Vishal Sunder, Brian Kingsbury, George Saon, Samuel Thomas, Slava Shechtman, Hagai Aronowitz, Eric Fosler-Lussier, Luis A. Lastras |
| 2025 | Interspeech | Exploring the Limits of Conformer CTC-Encoder for Speech Emotion Recognition using Large Language Models. | Edmilson da Silva Morais, Hagai Aronowitz, Aharon Satt, Ron Hoory, Avihu Dekel, Brian Kingsbury, George Saon |
| 2024 | ICASSP | Semi-Autoregressive Streaming ASR with Label Context. | Siddhant Arora, George Saon, Shinji Watanabe, Brian Kingsbury |
| 2024 | ICASSP | Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization. | A F M Saif, Xiaodong Cui, Han Shen, Songtao Lu, Brian Kingsbury, Tianyi Chen |
| 2024 | Interspeech | Exploring the limits of decoder-only models trained on public speech recognition corpora. | Ankit Gupta, George Saon, Brian Kingsbury |
| 2024 | Interspeech | M2ASR: Multilingual Multi-task Automatic Speech Recognition via Multi-objective Optimization. | A F M Saif, Lisha Chen, Xiaodong Cui, Songtao Lu, Brian Kingsbury, Tianyi Chen |
| 2023 | ICASSP | C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval. | Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas, Rogrio Feris, Brian Kingsbury, Leonid Karlinsky, David Harwath, Hilde Kuehne, James R. Glass |
| 2023 | ICASSP | Fine-Grained Textual Knowledge Transfer to Improve RNN Transducers for Speech Recognition and Understanding. | Vishal Sunder, Samuel Thomas, Hong-Kwang Jeff Kuo, Brian Kingsbury, Eric Fosler-Lussier |
| 2023 | ICASSP | Multi-Speaker Data Augmentation for Improved end-to-end Automatic Speech Recognition. | Samuel Thomas, Hong-Kwang Jeff Kuo, George Saon, Brian Kingsbury |
| 2023 | Interspeech | Improving RNN Transducer Acoustic Models for English Conversational Speech Recognition. | Xiaodong Cui, George Saon, Brian Kingsbury |
| 2023 | Interspeech | Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages. | Andrew Rouditchenko, Sameer Khurana, Samuel Thomas, Rogrio Feris, Leonid Karlinsky, Hilde Kuehne, David Harwath, Brian Kingsbury, James R. Glass |
| 2023 | Interspeech | ConvKT: Conversation-Level Knowledge Transfer for Context Aware End-to-End Spoken Language Understanding. | Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas, Hong-Kwang Jeff Kuo, Brian Kingsbury |
| 2023 | ISIT | High-Dimensional Smoothed Entropy Estimation via Dimensionality Reduction. | Kristjan H. Greenewald, Brian Kingsbury, Yuancheng Yu |
| 2022 | CVPR | Everything at Once - Multi-modal Fusion Transformer for Video Retrieval. | Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Hilde Kuehne |
| 2022 | ICASSP | A New Data Augmentation Method for Intent Classification Enhancement and its Application on Spoken Conversation Datasets. | Zvi Kons, Aharon Satt, Hong-Kwang Kuo, Samuel Thomas, Boaz Carmeli, Ron Hoory, Brian Kingsbury |
| 2022 | ICASSP | Improving End-to-end Models for Set Prediction in Spoken Language Understanding. | Hong-Kwang Jeff Kuo, Zoltn Tske, Samuel Thomas, Brian Kingsbury, George Saon |
| 2022 | ICASSP | Decentralized Bilevel Optimization for Personalized Client Learning. | Songtao Lu, Xiaodong Cui, Mark S. Squillante, Brian Kingsbury, Lior Horesh |
| 2022 | ICASSP | Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding. | Vishal Sunder, Samuel Thomas, Hong-Kwang Jeff Kuo, Jatin Ganhotra, Brian Kingsbury, Eric Fosler-Lussier |
| 2022 | ICASSP | Towards Reducing the Need for Speech Training Data to Build Spoken Language Understanding Systems. | Samuel Thomas, Hong-Kwang Jeff Kuo, Brian Kingsbury, George Saon |
| 2022 | ICASSP | Integrating Text Inputs for Training and Adapting RNN Transducer ASR Models. | Samuel Thomas, Brian Kingsbury, George Saon, Hong-Kwang Jeff Kuo |
| 2022 | Interspeech | Improving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label Smoothing. | Xiaodong Cui, George Saon, Tohru Nagano, Masayuki Suzuki, Takashi Fukuda, Brian Kingsbury, Gakuto Kurata |
| 2022 | Interspeech | Accelerating Inference and Language Model Fusion of Recurrent Neural Network Transducers via End-to-End 4-bit Quantization. | Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Swagath Venkataramani, George Saon, Xiaodong Cui, Brian Kingsbury, Kailash Gopalakrishnan |
| 2022 | Interspeech | Global RNN Transducer Models For Multi-dialect Speech Recognition. | Takashi Fukuda, Samuel Thomas, Masayuki Suzuki, Gakuto Kurata, George Saon, Brian Kingsbury |
| 2022 | Interspeech | VQ-T: RNN Transducers using Vector-Quantized Prediction Network States. | Jiatong Shi, George Saon, David Haws, Shinji Watanabe, Brian Kingsbury |
| 2022 | Interspeech | Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent Systems. | Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas, Hong-Kwang Kuo, Brian Kingsbury |
| 2021 | ICASSP | RNN Transducer Models for Spoken Language Understanding. | Samuel Thomas, Hong-Kwang Jeff Kuo, George Saon, Zoltn Tske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory |
| 2021 | ICASSP | Federated Acoustic Modeling for Automatic Speech Recognition. | Xiaodong Cui, Songtao Lu, Brian Kingsbury |
| 2021 | ICASSP | End-to-End Spoken Language Understanding Using Transformer Networks and Self-Supervised Pre-Trained Features. | Edmilson da Silva Morais, Hong-Kwang Jeff Kuo, Samuel Thomas, Zoltn Tske, Brian Kingsbury |
| 2021 | ICASSP | Advancing RNN Transducer Technology for Speech Recognition. | George Saon, Zoltn Tske, Daniel Bolaos, Brian Kingsbury |
| 2021 | ICCV | Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos. | Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang |
| 2021 | Interspeech | Reducing Exposure Bias in Training Recurrent Neural Network Transducers. | Xiaodong Cui, Brian Kingsbury, George Saon, David Haws, Zoltn Tske |
| 2021 | Interspeech | 4-Bit Quantization of LSTM-Based Speech Recognition Models. | Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Xiao Sun, Naigang Wang, Swagath Venkataramani, George Saon, Xiaodong Cui, Brian Kingsbury, Wei Zhang, Zoltn Tske, Kailash Gopalakrishnan |
| 2021 | Interspeech | Integrating Dialog History into End-to-End Spoken Language Understanding Systems. | Jatin Ganhotra, Samuel Thomas, Hong-Kwang Jeff Kuo, Sachindra Joshi, George Saon, Zoltn Tske, Brian Kingsbury |
| 2021 | Interspeech | Improving Customization of Neural Transducers by Mitigating Acoustic Mismatch of Synthesized Audio. | Gakuto Kurata, George Saon, Brian Kingsbury, David Haws, Zoltn Tske |
| 2021 | Interspeech | Cascaded Multilingual Audio-Visual Learning from Videos. | Andrew Rouditchenko, Angie W. Boggust, David Harwath, Samuel Thomas, Hilde Kuehne, Brian Chen, Rameswar Panda, Rogrio Feris, Brian Kingsbury, Michael Picheny, James R. Glass |
| 2021 | Interspeech | AVLnet: Learning Audio-Visual Language Representations from Instructional Videos. | Andrew Rouditchenko, Angie W. Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogrio Schmidt Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, James R. Glass |
| 2021 | Interspeech | On the Limit of English Conversational Speech Recognition. | Zoltn Tske, George Saon, Brian Kingsbury |
| 2020 | ICASSP | Fast Training of Deep Neural Networks for Speech Recognition. | Guojing Cong, Brian Kingsbury, Chih-Chieh Yang, Tianyi Liu |
| 2020 | ICASSP | Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent Systems. | Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas, Zvi Kons, Kartik Audhkhasi, Brian Kingsbury, Ron Hoory, Michael Picheny |
| 2020 | ICASSP | Improving Efficiency in Large-Scale Decentralized Distributed Training. | Wei Zhang, Xiaodong Cui, Abdullah Kayi, Mingrui Liu, Ulrich Finkler, Brian Kingsbury, George Saon, Youssef Mroueh, Alper Buyuktosunoglu, Payel Das, David S. Kung, Michael Picheny |
| 2020 | Interspeech | Transliteration Based Data Augmentation for Training Multilingual ASR Acoustic Models in Low Resource Settings. | Samuel Thomas, Kartik Audhkhasi, Brian Kingsbury |
| 2020 | Interspeech | End-to-End Spoken Language Understanding Without Full Transcripts. | Hong-Kwang Jeff Kuo, Zoltn Tske, Samuel Thomas, Yinghui Huang, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, Luis A. Lastras |
| 2020 | Interspeech | Representation Based Meta-Learning for Few-Shot Spoken Intent Recognition. | Ashish R. Mittal, Samarth Bharadwaj, Shreya Khare, Saneem A. Chemmengath, Karthik Sankaranarayanan, Brian Kingsbury |
| 2020 | Interspeech | Single Headed Attention Based Sequence-to-Sequence Model for State-of-the-Art Results on Switchboard. | Zoltn Tske, George Saon, Kartik Audhkhasi, Brian Kingsbury |
| 2019 | ASRU | Simplified LSTMS for Speech Recognition. | George Saon, Zoltn Tske, Kartik Audhkhasi, Brian Kingsbury, Michael Picheny, Samuel Thomas |
| 2019 | ICASSP | Sequence Noise Injected Training for End-to-end Speech Recognition. | George Saon, Zoltn Tske, Kartik Audhkhasi, Brian Kingsbury |
| 2019 | ICASSP | English Broadcast News Speech Recognition by Humans and Machines. | Samuel Thomas, Masayuki Suzuki, Yinghui Huang, Gakuto Kurata, Zoltn Tske, George Saon, Brian Kingsbury, Michael Picheny, Tom Dibert, Alice Kaiser-Schatzlein, Bern Samko |
| 2019 | ICASSP | Distributed Deep Learning Strategies for Automatic Speech Recognition. | Wei Zhang, Xiaodong Cui, Ulrich Finkler, Brian Kingsbury, George Saon, David S. Kung, Michael Picheny |
| 2019 | ICML | Beyond Backprop: Online Alternating Minimization with Auxiliary Variables. | Anna Choromanska, Benjamin Cowen, Sadhana Kumaravel, Ronny Luss, Mattia Rigotti, Irina Rish, Paolo Diachille, Viatcheslav Gurev, Brian Kingsbury, Ravi Tejwani, Djallel Bouneffouf |
| 2019 | ICML | Estimating Information Flow in Deep Neural Networks. | Ziv Goldfeld, Ewout van den Berg, Kristjan H. Greenewald, Igor Melnyk, Nam Nguyen, Brian Kingsbury, Yury Polyanskiy |
| 2019 | Interspeech | Forget a Bit to Learn Better: Soft Forgetting for CTC-Based Automatic Speech Recognition. | Kartik Audhkhasi, George Saon, Zoltn Tske, Brian Kingsbury, Michael Picheny |
| 2019 | Interspeech | Challenging the Boundaries of Speech Recognition: The MALACH Corpus. | Michael Picheny, Zoltn Tske, Brian Kingsbury, Kartik Audhkhasi, Xiaodong Cui, George Saon |
| 2019 | Interspeech | A Highly Efficient Distributed Deep Learning System for Automatic Speech Recognition. | Wei Zhang, Xiaodong Cui, Ulrich Finkler, George Saon, Abdullah Kayi, Alper Buyuktosunoglu, Brian Kingsbury, David S. Kung, Michael Picheny |
| 2018 | ICASSP | Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech Recognition. | Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Michael Picheny |
| 2017 | ICASSP | End-to-end ASR-free keyword search from speech. | Kartik Audhkhasi, Andrew Rosenberg, Abhinav Sethy, Bhuvana Ramabhadran, Brian Kingsbury |
| 2017 | ICASSP | Knowledge distillation across ensembles of multilingual models for low-resource languages. | Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Tom Sercu, Kartik Audhkhasi, Abhinav Sethy, Markus Nubaum-Thom, Andrew Rosenberg |
| 2017 | ICASSP | Network architectures for multilingual speech representation learning. | Tom Sercu, George Saon, Jia Cui, Xiaodong Cui, Bhuvana Ramabhadran, Brian Kingsbury, Abhinav Sethy |
| 2017 | SC | Accelerating deep neural network learning for speech recognition on a cluster of GPUs. | Guojing Cong, Brian Kingsbury, Soumyadip Gosh, George Saon, Fan Zhou |
| 2016 | ICASSP | Efficient one-vs-one kernel ridge regression for speech recognition. | Jie Chen, Lingfei Wu, Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran |
| 2016 | ICASSP | A comparison between deep neural nets and kernel acoustic models for speech recognition. | Zhiyun Lu, Dong Guo, Alireza Bagheri Garakani, Kuan Liu, Avner May, Aurlien Bellet, Linxi Fan, Michael Collins, Brian Kingsbury, Michael Picheny, Fei Sha |
| 2016 | ICASSP | Compact kernel models for acoustic modeling via random feature selection. | Avner May, Michael Collins, Daniel J. Hsu, Brian Kingsbury |
| 2016 | ICASSP | Very deep multilingual convolutional neural networks for LVCSR. | Tom Sercu, Christian Puhrsch, Brian Kingsbury, Yann LeCun |
| 2016 | Interspeech | Improved Neural Network Initialization by Grouping Context-Dependent Targets for Acoustic Modeling. | Gakuto Kurata, Brian Kingsbury |
| 2016 | Interspeech | Multilingual Data Selection for Low Resource Speech Recognition. | Samuel Thomas, Kartik Audhkhasi, Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran |
| 2015 | ASRU | Multilingual representations for low resource speech recognition and keyword search. | Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Xiaodong Cui, Ellen Kislal, Lidia Mangu, Markus Nubaum-Thom, Michael Picheny, Zoltn Tske, Pavel Golik, Ralf Schlter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Haipeng Wang, Philip C. Woodland |
| 2015 | ICASSP | Data augmentation for deep convolutional neural network acoustic modeling. | Xiaodong Cui, Vaibhava Goel, Brian Kingsbury |
| 2015 | ICASSP | Order-free spoken term detection. | Lidia Mangu, George Saon, Michael Picheny, Brian Kingsbury |
| 2015 | Interspeech | A multi-region deep neural network model in speech recognition. | Jia Cui, George Saon, Bhuvana Ramabhadran, Brian Kingsbury |
| 2014 | ICASSP | Data Augmentation for deep neural network acoustic modeling. | Xiaodong Cui, Vaibhava Goel, Brian Kingsbury |
| 2014 | ICASSP | Automatic keyword selection for keyword search development and tuning. | Jia Cui, Jonathan Mamou, Brian Kingsbury, Bhuvana Ramabhadran |
| 2014 | ICASSP | Efficient spoken term detection using confusion networks. | Lidia Mangu, Brian Kingsbury, Hagen Soltau, Hong-Kwang Kuo, Michael Picheny |
| 2014 | ICASSP | Improvements to filterbank and delta learning within a deep neural network framework. | Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George Saon, Bhuvana Ramabhadran |
| 2014 | Interspeech | Improving deep neural network acoustic modeling for audio corpus indexing under the IARPA babel program. | Xiaodong Cui, Brian Kingsbury, Jia Cui, Bhuvana Ramabhadran, Andrew Rosenberg, Mohammad Sadegh Rasooli, Owen Rambow, Nizar Habash, Vaibhava Goel |
| 2014 | Interspeech | Recent improvements in neural network acoustic modeling for LVCSR in low resource languages. | Jia Cui, Bhuvana Ramabhadran, Xiaodong Cui, Andrew Rosenberg, Brian Kingsbury, Abhinav Sethy |
| 2014 | Interspeech | Parallel deep neural network training for LVCSR tasks using blue gene/Q. | Tara N. Sainath, I-Hsin Chung, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Brian Kingsbury, George Saon, Vernon Austel, Upendra V. Chaudhari |
| 2014 | Interspeech | Deep scattering spectra with deep neural networks for LVCSR tasks. | Tara N. Sainath, Vijayaditya Peddinti, Brian Kingsbury, Petr Fousek, Bhuvana Ramabhadran, David Nahamoo |
| 2014 | SC | Parallel Deep Neural Network Training for Big Data on Blue Gene/Q. | I-Hsin Chung, Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Vernon Austel, Upendra V. Chaudhari, Brian Kingsbury |
| 2013 | ASRU | Accelerating Hessian-free optimization for Deep Neural Networks by implicit preconditioning and sampling. | Tara N. Sainath, Lior Horesh, Brian Kingsbury, Aleksandr Y. Aravkin, Bhuvana Ramabhadran |
| 2013 | ASRU | Improvements to Deep Convolutional Neural Networks for LVCSR. | Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George E. Dahl, George Saon, Hagen Soltau, Toms Beran, Aleksandr Y. Aravkin, Bhuvana Ramabhadran |
| 2013 | ASRU | Learning filter banks within a deep neural network framework. | Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, Bhuvana Ramabhadran |
| 2013 | ASRU | An empirical study of confusion modeling in keyword search for low resource languages. | Murat Saraclar, Abhinav Sethy, Bhuvana Ramabhadran, Lidia Mangu, Jia Cui, Xiaodong Cui, Brian Kingsbury, Jonathan Mamou |
| 2013 | ICASSP | Developing speech recognition systems for corpus indexing under the IARPA Babel program. | Jia Cui, Xiaodong Cui, Bhuvana Ramabhadran, Janice Kim, Brian Kingsbury, Jonathan Mamou, Lidia Mangu, Michael Picheny, Tara N. Sainath, Abhinav Sethy |
| 2013 | ICASSP | New types of deep neural network learning for speech recognition and related applications: an overview. | Li Deng, Geoffrey E. Hinton, Brian Kingsbury |
| 2013 | ICASSP | Audio-visual deep learning for noise robust speech recognition. | Jing Huang, Brian Kingsbury |
| 2013 | ICASSP | A high-performance Cantonese keyword search system. | Brian Kingsbury, Jia Cui, Xiaodong Cui, Mark J. F. Gales, Kate M. Knill, Jonathan Mamou, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlter, Abhinav Sethy, Philip C. Woodland |
| 2013 | ICASSP | System combination and score normalization for spoken term detection. | Jonathan Mamou, Jia Cui, Xiaodong Cui, Mark J. F. Gales, Brian Kingsbury, Kate M. Knill, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlter, Abhinav Sethy, Philip C. Woodland |
| 2013 | ICASSP | Exploiting diversity for spoken term detection. | Lidia Mangu, Hagen Soltau, Hong-Kwang Kuo, Brian Kingsbury, George Saon |
| 2013 | ICASSP | Low-rank matrix factorization for Deep Neural Network training with high-dimensional output targets. | Tara N. Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, Bhuvana Ramabhadran |
| 2013 | ICASSP | Deep convolutional neural networks for LVCSR. | Tara N. Sainath, Abdel-rahman Mohamed, Brian Kingsbury, Bhuvana Ramabhadran |
| 2013 | Interspeech | Mixtures of Bayesian joint factor analyzers for noise robust automatic speech recognition. | Xiaodong Cui, Vaibhava Goel, Brian Kingsbury |
| 2013 | Interspeech | The IBM speech activity detection system for the DARPA RATS program. | George Saon, Samuel Thomas, Hagen Soltau, Sriram Ganapathy, Brian Kingsbury |
| 2012 | ICASSP | Auto-encoder bottleneck features using deep belief networks. | Tara N. Sainath, Brian Kingsbury, Bhuvana Ramabhadran |
| 2012 | Interspeech | Scalable Minimum Bayes Risk Training of Deep Neural Network Acoustic Models Using Distributed Hessian-free Optimization. | Brian Kingsbury, Tara N. Sainath, Hagen Soltau |
| 2012 | Interspeech | Discriminative feature-space transforms using deep neural networks. | George Saon, Brian Kingsbury |
| 2012 | NAACL | Deep Neural Network Language Models. | Ebru Arisoy, Tara N. Sainath, Brian Kingsbury, Bhuvana Ramabhadran |
| 2011 | ASRU | The IBM 2011 GALE Arabic speech transcription system. | Lidia Mangu, Hong-Kwang Kuo, Stephen M. Chu, Brian Kingsbury, George Saon, Hagen Soltau, Fadi Biadsy |
| 2011 | ASRU | Making Deep Belief Networks effective for large vocabulary continuous speech recognition. | Tara N. Sainath, Brian Kingsbury, Bhuvana Ramabhadran, Petr Fousek, Petr Novk, Abdel-rahman Mohamed |
| 2011 | ICASSP | Arccosine kernels: Acoustic modeling with infinite neural networks. | Chih-Chieh Cheng, Brian Kingsbury |
| 2011 | ICASSP | The IBM 2009 GALE Arabic speech transcription system. | Brian Kingsbury, Hagen Soltau, George Saon, Stephen M. Chu, Hong-Kwang Kuo, Lidia Mangu, Suman V. Ravuri, Nelson Morgan, Adam Janin |
| 2010 | ICASSP | The IBM 2008 GALE Arabic speech transcription system. | George Saon, Hagen Soltau, Upendra V. Chaudhari, Stephen M. Chu, Brian Kingsbury, Hong-Kwang Kuo, Lidia Mangu, Daniel Povey |
| 2009 | ICASSP | Lattice-based optimization of sequence classification criteria for neural-network acoustic modeling. | Brian Kingsbury |
| 2009 | NAACL | Fast decoding for open vocabulary spoken term detection. | Bhuvana Ramabhadran, Abhinav Sethy, Jonathan Mamou, Brian Kingsbury, Upendra V. Chaudhari |
| 2009 | NAACL | Tied-Mixture Language Modeling in Continuous Space. | Ruhi Sarikaya, Mohamed Afify, Brian Kingsbury |
| 2008 | ICASSP | Boosted MMI for model and feature-space discriminative training. | Daniel Povey, Dimitri Kanevsky, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Karthik Visweswariah |
| 2008 | Interspeech | Discriminative graph training for ultra-fast low-footprint speech indexing. | Upendra V. Chaudhari, Hong-Kwang Jeff Kuo, Brian Kingsbury |
| 2008 | Interspeech | Monte Carlo model-space noise adaptation for speech recognition. | Daniel Povey, Brian Kingsbury |
| 2008 | Interspeech | Machine translation in continuous space. | Ruhi Sarikaya, Yonggang Deng, Mohamed Afify, Brian Kingsbury, Yuqing Gao |
| 2007 | ICASSP | Discriminative Training of Decoding Graphs for Large Vocabulary Continuous Speech Recognition. | Hong-Kwang Jeff Kuo, Brian Kingsbury, Geoffrey Zweig |
| 2007 | ICASSP | Evaluation of Proposed Modifications to MPE for Large Scale Discriminative Training. | Daniel Povey, Brian Kingsbury |
| 2007 | ICASSP | The IBM 2006 Gale Arabic ASR System. | Hagen Soltau, George Saon, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey, Geoffrey Zweig |
| 2006 | ICASSP | Automated Quality Monitoring in the Call Center with ASR and Maximum Entropy. | Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
| 2006 | NAACL | Automated Quality Monitoring for Call Centers using Speech and NLP Technologies. | Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
| 2005 | ICASSP | fMPE: Discriminatively Trained Features for Speech Recognition. | Daniel Povey, Brian Kingsbury, Lidia Mangu, George Saon, Hagen Soltau, Geoffrey Zweig |
| 2005 | ICASSP | Contructing Ensembles of ASR Systems Using Randomized Decision Trees. | Olivier Siohan, Bhuvana Ramabhadran, Brian Kingsbury |
| 2005 | ICASSP | The IBM 2004 Conversational Telephony System for Rich Transcription. | Hagen Soltau, Brian Kingsbury, Lidia Mangu, Daniel Povey, George Saon, Geoffrey Zweig |
| 2004 | ICASSP | An evaluation of a nonlinear feature transformation for conversational speech recognition. | Mohamed K. Omar, Brian Kingsbury |
| 2003 | Interspeech | Large vocabulary conversational speech recognition with a subspace constraint on inverse covariance matrices. | Scott Axelrod, Vaibhava Goel, Brian Kingsbury, Karthik Visweswariah, Ramesh A. Gopinath |
| 2003 | Interspeech | Toward domain-independent conversational speech recognition. | Brian Kingsbury, Lidia Mangu, George Saon, Geoffrey Zweig, Scott Axelrod, Vaibhava Goel, Karthik Visweswariah, Michael Picheny |
| 2003 | Interspeech | An architecture for rapid decoding of large vocabulary conversational speech. | George Saon, Geoffrey Zweig, Brian Kingsbury, Lidia Mangu, Upendra V. Chaudhari |
| 2002 | ICASSP | Robust speech recognition in Noisy Environments: The 2001 IBM spine evaluation system. | Brian Kingsbury, George Saon, Lidia Mangu, Mukund Padmanabhan, Ruhi Sarikaya |
| 2002 | Interspeech | Large vocabulary conversational speech recognition with the extended maximum likelihood linear transformation (EMLLT) model. | Jing Huang, Vaibhava Goel, Ramesh Gopinath, Brian Kingsbury, Peder A. Olsen, Karthik Visweswariah |
| 2002 | Interspeech | Distributed speech recognition using noise-robust MFCC and traps-estimated manner features. | Pratibha Jain, Hynek Hermansky, Brian Kingsbury |
| 2002 | Interspeech | A hybrid HMM/traps model for robust voice activity detection. | Brian Kingsbury, Pratibha Jain, Andr Gustavo Adami |
| 2000 | Interspeech | Recent improvements in speech recognition performance on large vocabulary conversational speech (voicemail and switchboard). | Jing Huang, Brian Kingsbury, Lidia Mangu, Mukund Padmanabhan, George Saon, Geoffrey Zweig |
| 1998 | ICASSP | Incorporating information from syllable-length time scales into automatic speech recognition. | Su-Lin Wu, Brian Kingsbury, Nelson Morgan, Steven Greenberg |
| 1998 | Interspeech | Performance improvements through combining phone- and syllable-scale information in automatic speech recognition. | Su-Lin Wu, Brian Kingsbury, Nelson Morgan, Steven Greenberg |
| 1997 | ICASSP | The modulation spectrogram: in pursuit of an invariant representation of speech. | Steven Greenberg, Brian Kingsbury |
| 1997 | ICASSP | Recognizing reverberant speech with RASTA-PLP. | Brian Kingsbury, Nelson Morgan |