| 2025 | ASRU | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities. | George Saon, Avihu Dekel, Alexander Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Xiaodong Cui, Zvi Kons |
| 2025 | ICLR | Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis. | Hongkang Li, Songtao Lu, Pin-Yu Chen, Xiaodong Cui, Meng Wang |
| 2025 | SMC | SEER: Semantic Enhancement and Emotional Reasoning Network for Multimodal Fake News Detection. | Peican Zhu, Yubo Jing, Le Cheng, Bin Chen, Xiaodong Cui, Lianwei Wu, Keke Tang |
| 2024 | ICASSP | Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization. | A F M Saif, Xiaodong Cui, Han Shen, Songtao Lu, Brian Kingsbury, Tianyi Chen |
| 2024 | ICASSP | Reparameterization Head for Efficient Multi-Input Networks. | Keke Tang, Wenyu Zhao, Weilong Peng, Xiang Fang, Xiaodong Cui, Peican Zhu, Zhihong Tian |
| 2024 | ICASSP | How Can Personalized Context Help? Exploring Joint Retrieval of Passage and Personalized Context. | Hui Wan, Hongkang Li, Songtao Lu, Xiaodong Cui, Marina Danilevsky |
| 2024 | ICML | How Do Nonlinear Transformers Learn and Generalize in In-Context Learning? | Hongkang Li, Meng Wang, Songtao Lu, Xiaodong Cui, Pin-Yu Chen |
| 2024 | Interspeech | M2ASR: Multilingual Multi-task Automatic Speech Recognition via Multi-objective Optimization. | A F M Saif, Lisha Chen, Xiaodong Cui, Songtao Lu, Brian Kingsbury, Tianyi Chen |
| 2023 | AISTATS | Distributed Offline Policy Optimization Over Batch Data. | Han Shen, Songtao Lu, Xiaodong Cui, Tianyi Chen |
| 2023 | CIKM | HEPT Attack: Heuristic Perpendicular Trial for Hard-label Attacks under Limited Query Budgets. | Qi Li, Xingyu Li, Xiaodong Cui, Keke Tang, Peican Zhu |
| 2023 | ICASSP | Diagonal State Space Augmented Transformers for Speech Recognition. | George Saon, Ankit Gupta, Xiaodong Cui |
| 2023 | ICML | Compressed Decentralized Proximal Stochastic Gradient Method for Nonconvex Composite Problems with Heterogeneous Data. | Yonggui Yan, Jie Chen, Pin-Yu Chen, Xiaodong Cui, Songtao Lu, Yangyang Xu |
| 2023 | Interspeech | Improving RNN Transducer Acoustic Models for English Conversational Speech Recognition. | Xiaodong Cui, George Saon, Brian Kingsbury |
| 2022 | ICASSP | Decentralized Bilevel Optimization for Personalized Client Learning. | Songtao Lu, Xiaodong Cui, Mark S. Squillante, Brian Kingsbury, Lior Horesh |
| 2022 | Interspeech | Improving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label Smoothing. | Xiaodong Cui, George Saon, Tohru Nagano, Masayuki Suzuki, Takashi Fukuda, Brian Kingsbury, Gakuto Kurata |
| 2022 | Interspeech | Accelerating Inference and Language Model Fusion of Recurrent Neural Network Transducers via End-to-End 4-bit Quantization. | Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Swagath Venkataramani, George Saon, Xiaodong Cui, Brian Kingsbury, Kailash Gopalakrishnan |
| 2021 | ACL | On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation. | Wei Zhang, Ziming Huang, Yada Zhu, Guangnan Ye, Xiaodong Cui, Fan Zhang |
| 2021 | ICASSP | Federated Acoustic Modeling for Automatic Speech Recognition. | Xiaodong Cui, Songtao Lu, Brian Kingsbury |
| 2021 | ICASSP | Speech Emotion Recognition with Multiscale Area Attention and Data Augmentation. | Mingke Xu, Fan Zhang, Xiaodong Cui, Wei Zhang |
| 2021 | Interspeech | Reducing Exposure Bias in Training Recurrent Neural Network Transducers. | Xiaodong Cui, Brian Kingsbury, George Saon, David Haws, Zoltn Tske |
| 2021 | Interspeech | 4-Bit Quantization of LSTM-Based Speech Recognition Models. | Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Xiao Sun, Naigang Wang, Swagath Venkataramani, George Saon, Xiaodong Cui, Brian Kingsbury, Wei Zhang, Zoltn Tske, Kailash Gopalakrishnan |
| 2020 | ICASSP | Improving Efficiency in Large-Scale Decentralized Distributed Training. | Wei Zhang, Xiaodong Cui, Abdullah Kayi, Mingrui Liu, Ulrich Finkler, Brian Kingsbury, George Saon, Youssef Mroueh, Alper Buyuktosunoglu, Payel Das, David S. Kung, Michael Picheny |
| 2020 | ICLR | Towards Better Understanding of Adaptive Gradient Algorithms in Generative Adversarial Nets. | Mingrui Liu, Youssef Mroueh, Jerret Ross, Wei Zhang, Xiaodong Cui, Payel Das, Tianbao Yang |
| 2020 | IJCAI | Task-Based Learning via Task-Oriented Prediction Network with Applications in Finance. | Di Chen, Yada Zhu, Xiaodong Cui, Carla P. Gomes |
| 2020 | KDD | Map Generation from Large Scale Incomplete and Inaccurate Data Labels. | Rui Zhang, Conrad M. Albrecht, Wei Zhang, Xiaodong Cui, Ulrich Finkler, David S. Kung, Siyuan Lu |
| 2019 | ICASSP | Cyclegan Bandwidth Extension Acoustic Modeling for Automatic Speech Recognition. | David Haws, Xiaodong Cui |
| 2019 | ICASSP | Distributed Deep Learning Strategies for Automatic Speech Recognition. | Wei Zhang, Xiaodong Cui, Ulrich Finkler, Brian Kingsbury, George Saon, David S. Kung, Michael Picheny |
| 2019 | Interspeech | Acoustic Model Optimization Based on Evolutionary Stochastic Gradient Descent with Anchors for Automatic Speech Recognition. | Xiaodong Cui, Michael Picheny |
| 2019 | Interspeech | Large-Scale Mixed-Bandwidth Deep Neural Network Acoustic Modeling for Automatic Speech Recognition. | Khoi-Nguyen C. Mac, Xiaodong Cui, Wei Zhang, Michael Picheny |
| 2019 | Interspeech | Challenging the Boundaries of Speech Recognition: The MALACH Corpus. | Michael Picheny, Zoltn Tske, Brian Kingsbury, Kartik Audhkhasi, Xiaodong Cui, George Saon |
| 2019 | Interspeech | A Highly Efficient Distributed Deep Learning System for Automatic Speech Recognition. | Wei Zhang, Xiaodong Cui, Ulrich Finkler, George Saon, Abdullah Kayi, Alper Buyuktosunoglu, Brian Kingsbury, David S. Kung, Michael Picheny |
| 2017 | ICASSP | Network architectures for multilingual speech representation learning. | Tom Sercu, George Saon, Jia Cui, Xiaodong Cui, Bhuvana Ramabhadran, Brian Kingsbury, Abhinav Sethy |
| 2017 | Interspeech | Embedding-Based Speaker Adaptive Training of Deep Neural Networks. | Xiaodong Cui, Vaibhava Goel, George Saon |
| 2017 | Interspeech | English Conversational Telephone Speech Recognition by Humans and Machines. | George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas, Dimitrios Dimitriadis, Xiaodong Cui, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
| 2016 | ICASSP | Efficient non-linear feature adaptation using Maxout networks. | Steven J. Rennie, Xiaodong Cui, Vaibhava Goel |
| 2015 | ASRU | Multilingual representations for low resource speech recognition and keyword search. | Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Xiaodong Cui, Ellen Kislal, Lidia Mangu, Markus Nubaum-Thom, Michael Picheny, Zoltn Tske, Pavel Golik, Ralf Schlter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Haipeng Wang, Philip C. Woodland |
| 2015 | ICASSP | Maximum likelihood nonlinear transformations based on deep neural networks. | Xiaodong Cui, Vaibhava Goel |
| 2015 | ICASSP | Data augmentation for deep convolutional neural network acoustic modeling. | Xiaodong Cui, Vaibhava Goel, Brian Kingsbury |
| 2015 | ICASSP | Annealed dropout trained maxout networks for improved LVCSR. | Steven J. Rennie, Pierre L. Dognin, Xiaodong Cui, Vaibhava Goel |
| 2014 | ICASSP | Data Augmentation for deep neural network acoustic modeling. | Xiaodong Cui, Vaibhava Goel, Brian Kingsbury |
| 2014 | ICASSP | A family of discriminative training criteria based on the F-divergence for deep neural networks. | Markus Nubaum-Thom, Xiaodong Cui, Ralf Schlter, Vaibhava Goel, Hermann Ney |
| 2014 | Interspeech | Improving deep neural network acoustic modeling for audio corpus indexing under the IARPA babel program. | Xiaodong Cui, Brian Kingsbury, Jia Cui, Bhuvana Ramabhadran, Andrew Rosenberg, Mohammad Sadegh Rasooli, Owen Rambow, Nizar Habash, Vaibhava Goel |
| 2014 | Interspeech | Recent improvements in neural network acoustic modeling for LVCSR in low resource languages. | Jia Cui, Bhuvana Ramabhadran, Xiaodong Cui, Andrew Rosenberg, Brian Kingsbury, Abhinav Sethy |
| 2014 | Interspeech | Exploiting vocal-source features to improve ASR accuracy for low-resource languages. | Raul Fernandez, Jia Cui, Andrew Rosenberg, Bhuvana Ramabhadran, Xiaodong Cui |
| 2013 | ASRU | An empirical study of confusion modeling in keyword search for low resource languages. | Murat Saraclar, Abhinav Sethy, Bhuvana Ramabhadran, Lidia Mangu, Jia Cui, Xiaodong Cui, Brian Kingsbury, Jonathan Mamou |
| 2013 | ICASSP | Developing speech recognition systems for corpus indexing under the IARPA Babel program. | Jia Cui, Xiaodong Cui, Bhuvana Ramabhadran, Janice Kim, Brian Kingsbury, Jonathan Mamou, Lidia Mangu, Michael Picheny, Tara N. Sainath, Abhinav Sethy |
| 2013 | ICASSP | A high-performance Cantonese keyword search system. | Brian Kingsbury, Jia Cui, Xiaodong Cui, Mark J. F. Gales, Kate M. Knill, Jonathan Mamou, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlter, Abhinav Sethy, Philip C. Woodland |
| 2013 | ICASSP | System combination and score normalization for spoken term detection. | Jonathan Mamou, Jia Cui, Xiaodong Cui, Mark J. F. Gales, Brian Kingsbury, Kate M. Knill, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlter, Abhinav Sethy, Philip C. Woodland |
| 2013 | Interspeech | Mixtures of Bayesian joint factor analyzers for noise robust automatic speech recognition. | Xiaodong Cui, Vaibhava Goel, Brian Kingsbury |
| 2013 | Interspeech | Adaptive stereo-based stochastic mapping. | Shay Maymon, Pierre L. Dognin, Xiaodong Cui, Vaibhava Goel |
| 2012 | ICASSP | Stereo-based stochastic mapping with context using probabilistic PCA for noise robust automatic speech recognition. | Xiaodong Cui, Mohamed Afify, Bowen Zhou |
| 2012 | Interspeech | Sparse Bayesian Factor Analysis for Stereo-based Stochastic Mapping. | Xiaodong Cui, Mohamed Afify, George Saon, Vaibhava Goel |
| 2011 | ASRU | An investigation of heuristic, manual and statistical pronunciation derivation for Pashto. | Upendra V. Chaudhari, Xiaodong Cui, Bowen Zhou, Rong Zhang |
| 2011 | ICASSP | Clustering of bootstrapped acoustic model with full covariance. | Xin Chen, Xiaodong Cui, Jian Xue, Peder A. Olsen, John R. Hershey, Bowen Zhou, Yunxin Zhao |
| 2011 | ICASSP | Multi-view and multi-objective semi-supervised learning for large vocabulary continuous speech recognition. | Xiaodong Cui, Jing Huang, Jen-Tzung Chien |
| 2011 | Interspeech | Acoustic Modeling with Bootstrap and Restructuring Based on Full Covariance. | Xiaodong Cui, Xin Chen, Jian Xue, Peder A. Olsen, John R. Hershey, Bowen Zhou |
| 2011 | Interspeech | Towards High Performance LVCSR in Speech-to-Speech Translation System on Smart Phones. | Jian Xue, Xiaodong Cui, Gregg Daggett, Etienne Marcheret, Bowen Zhou |
| 2010 | ICASSP | A comparative study on system combination schemes for LVCSR. | Chengyuan Ma, Hong-Kwang Jeff Kuo, Hagen Soltau, Xiaodong Cui, Upendra V. Chaudhari, Lidia Mangu, Chin-Hui Lee |
| 2010 | Interspeech | Acoustic modeling with bootstrap and restructuring for low-resourced languages. | Xiaodong Cui, Jian Xue, Pierre L. Dognin, Upendra V. Chaudhari, Bowen Zhou |
| 2010 | Interspeech | Applying scalable phonetic context similarity in unit selection of concatenative text-to-speech. | Wei Zhang, Xiaodong Cui |
| 2009 | ASRU | Improving online incremental speaker adaptation with eigen feature space MLLR. | Xiaodong Cui, Jian Xue, Bowen Zhou |
| 2009 | ICASSP | Stereo-based stochastic mapping with discriminative training for noise robust speech recognition. | Xiaodong Cui, Mohamed Afify, Yuqing Gao |
| 2009 | Interspeech | A study of bootstrapping with multiple acoustic features for improved automatic speech recognition. | Xiaodong Cui, Jian Xue, Bing Xiang, Bowen Zhou |
| 2008 | ICASSP | MMSE-based stereo feature stochastic mapping for noise robust speech recognition. | Xiaodong Cui, Mohamed Afify, Yuqing Gao |
| 2008 | ICASSP | Developing high performance asr in the IBM multilingual speech-to-speech translation system. | Xiaodong Cui, Liang Gu, Bing Xiang, Wei Zhang, Yuqing Gao |
| 2008 | Interspeech | N-best based stochastic mapping on stereo HMM for noise robust speech recognition. | Xiaodong Cui, Mohamed Afify, Yuqing Gao |
| 2008 | Interspeech | High-performance low-latency speech recognition via multi-layered feature streaming and fast Gaussian computation. | Liang Gu, Jian Xue, Xiaodong Cui, Yuqing Gao |
| 2007 | ICASSP | Stereo-Based Stochastic Mapping for Robust Speech Recognition. | Mohamed Afify, Xiaodong Cui, Yuqing Gao |
| 2006 | ICASSP | Modeling Variance Variation in a Variable Parameter HMM Framework for Noise Robust Speech Recognition. | Xiaodong Cui, Yifan Gong |
| 2006 | ICASSP | A Database of Vocal Tract Resonance Trajectories for Research in Speech Processing. | Li Deng, Xiaodong Cui, Robert Pruvenok, Yanyi Chen, Safiyy Momen, Abeer Alwan |
| 2006 | Interspeech | Rapid speaker adaptation using regression-tree based spectral peak alignment. | Shizhen Wang, Xiaodong Cui, Abeer Alwan |
| 2005 | Interspeech | MLLR-like speaker adaptation based on linearization of VTLN with MFCC features. | Xiaodong Cui, Abeer Alwan |
| 2005 | Interspeech | TBALL data collection: the making of a young children's speech corpus. | Abe Kazemzadeh, Hong You, Markus Iseli, Barbara Jones, Xiaodong Cui, Margaret Heritage, Patti Price, Elaine Andersen, Shrikanth S. Narayanan, Abeer Alwan |
| 2004 | ICASSP | Can back-ends be more robust than front-ends? Investigation over the Aurora-2 database. | Alexis Bernard, Yifan Gong, Xiaodong Cui |
| 2004 | ICASSP | Combining feature compensation and weighted Viterbi decoding for noise robust speech recognition with limited adaptation data. | Xiaodong Cui, Abeer Alwan |
| 2003 | ICASSP | Variable parameter Gaussian mixture hidden Markov modeling for speech recognition. | Xiaodong Cui, Yifan Gong |
| 2003 | Interspeech | A noise-robust ASR back-end technique based on weighted viterbi recognition. | Xiaodong Cui, Alexis Bernard, Abeer Alwan |
| 2002 | ICASSP | Efficient adaptation text design based on the Kullback-Leibler measure. | Xiaodong Cui, Abeer Alwan |
| 2002 | Interspeech | Evaluation of noise robust features on the Aurora databases. | Xiaodong Cui, Markus Iseli, Qifeng Zhu, Abeer Alwan |
| 2001 | Interspeech | Noise robust feature extraction for ASR using the Aurora 2 database. | Qifeng Zhu, Markus Iseli, Xiaodong Cui, Abeer Alwan |
| 2000 | Interspeech | A language model adaptation approach based on text classification. | Jiasong Sun, Xiaodong Cui, Zuoying Wang, Yang Liu |