| 2024 | ICASSP | Improving Speech Recognition for African American English with Audio Classification. | Shefali Garg, Zhouyuan Huo, Khe Chai Sim, Suzan Schwartz, Mason Chua, Alna Aksnova, Tsendsuren Munkhdalai, Levi King, Darryl Wright, Zion Mengesha, Dongseong Hwang, Tara N. Sainath, Franoise Beaufays, Pedro Moreno Mengibar |
| 2024 | ICASSP | A Comparison of Parameter-Efficient ASR Domain Adaptation Methods for Universal Speech and Language Models. | Khe Chai Sim, Zhouyuan Huo, Tsendsuren Munkhdalai, Nikhil Siddhartha, Adam Stooke, Zhong Meng, Bo Li, Tara N. Sainath |
| 2024 | Interspeech | AdaRA: Adaptive Rank Allocation of Residual Adapters for Speech Foundation Model. | Zhouyuan Huo, Dongseong Hwang, Gan Song, Khe Chai Sim, Weiran Wang |
| 2024 | Interspeech | Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm. | Weiran Wang, Zelin Wu, Diamantino Caseiro, Tsendsuren Munkhdalai, Khe Chai Sim, Pat Rondon, Golan Pundak, Gan Song, Rohit Prabhavalkar, Zhong Meng, Ding Zhao, Tara Sainath, Yanzhang He, Pedro Moreno Mengibar |
| 2024 | NAACL | Massive End-to-end Speech Recognition Models with Time Reduction. | Weiran Wang, Rohit Prabhavalkar, Haozhe Shan, Zhong Meng, Dongseong Hwang, Qiujia Li, Khe Chai Sim, Bo Li, James Qin, Xingyu Cai, Adam Stooke, Chengjian Zheng, Yanzhang He, Tara N. Sainath, Pedro Moreno Mengibar |
| 2023 | ASRU | Contextual Spelling Correction with Large Language Models. | Gan Song, Zelin Wu, Golan Pundak, Angad Chandorkar, Kandarp Joshi, Xavier Velez, Diamantino Caseiro, Ben Haynor, Weiran Wang, Nikhil Siddhartha, Pat Rondon, Khe Chai Sim |
| 2023 | ICASSP | Resource-Efficient Transfer Learning from Speech Foundation Model Using Hierarchical Feature Fusion. | Zhouyuan Huo, Khe Chai Sim, Bo Li, Dongseong Hwang, Tara N. Sainath, Trevor Strohman |
| 2023 | ICASSP | Comparison of Soft and Hard Target RNN-T Distillation for Large-Scale ASR. | Dongseong Hwang, Khe Chai Sim, Yu Zhang, Trevor Strohman |
| 2023 | ICASSP | Efficient Domain Adaptation for Speech Foundation Models. | Bo Li, Dongseong Hwang, Zhouyuan Huo, Junwen Bai, Guru Prakash, Tara N. Sainath, Khe Chai Sim, Yu Zhang, Wei Han, Trevor Strohman, Franoise Beaufays |
| 2023 | Interspeech | Re-investigating the Efficient Transfer Learning of Speech Foundation Model using Feature Fusion Methods. | Zhouyuan Huo, Khe Chai Sim, Dongseong Hwang, Tsendsuren Munkhdalai, Tara N. Sainath, Pedro Moreno Mengibar |
| 2023 | Interspeech | Dual-Mode NAM: Effective Top-K Context Injection for End-to-End ASR. | Zelin Wu, Tsendsuren Munkhdalai, Pat Rondon, Golan Pundak, Khe Chai Sim, Christopher Li |
| 2022 | ICASSP | Joint Unsupervised and Supervised Training for Multilingual ASR. | Junwen Bai, Bo Li, Yu Zhang, Ankur Bapna, Nikhil Siddhartha, Khe Chai Sim, Tara N. Sainath |
| 2022 | ICASSP | Large-Scale ASR Domain Adaptation Using Self- and Semi-Supervised Learning. | Dongseong Hwang, Ananya Misra, Zhouyuan Huo, Nikhil Siddhartha, Shefali Garg, David Qiu, Khe Chai Sim, Trevor Strohman, Franoise Beaufays, Yanzhang He |
| 2022 | ICASSP | Fast Contextual Adaptation with Neural Associative Memory for On-Device Personalized Speech Recognition. | Tsendsuren Munkhdalai, Khe Chai Sim, Angad Chandorkar, Fan Gao, Mason Chua, Trevor Strohman, Franoise Beaufays |
| 2022 | Interspeech | UserLibri: A Dataset for ASR Personalization Using Only Text. | Theresa Breiner, Swaroop Ramaswamy, Ehsan Variani, Shefali Garg, Rajiv Mathews, Khe Chai Sim, Kilol Gupta, Mingqing Chen, Lara McConnaughey |
| 2022 | Interspeech | Incremental Layer-Wise Self-Supervised Learning for Efficient Unsupervised Speech Domain Adaptation On Device. | Zhouyuan Huo, Dongseong Hwang, Khe Chai Sim, Shefali Garg, Ananya Misra, Nikhil Siddhartha, Trevor Strohman, Franoise Beaufays |
| 2022 | Interspeech | Pseudo Label Is Better Than Human Label. | Dongseong Hwang, Khe Chai Sim, Zhouyuan Huo, Trevor Strohman |
| 2022 | Interspeech | On-the-fly ASR Corrections with Audio Exemplars. | Golan Pundak, Tsendsuren Munkhdalai, Khe Chai Sim |
| 2021 | Interspeech | A Comparison of Supervised and Unsupervised Pre-Training of End-to-End Models. | Ananya Misra, Dongseong Hwang, Zhouyuan Huo, Shefali Garg, Nikhil Siddhartha, Arun Narayanan, Khe Chai Sim |
| 2021 | Interspeech | Robust Continuous On-Device Personalization for Automatic Speech Recognition. | Khe Chai Sim, Angad Chandorkar, Fan Gao, Mason Chua, Tsendsuren Munkhdalai, Franoise Beaufays |
| 2020 | ICASSP | Low-Rank Gradient Approximation for Memory-Efficient on-Device Training of Deep Neural Network. | Mary Gooneratne, Khe Chai Sim, Petr Zadrazil, Andreas Kabel, Franoise Beaufays, Giovanni Motta |
| 2019 | ASRU | Personalization of End-to-End Speech Recognition on Mobile Devices for Named Entities. | Khe Chai Sim, Leif Johnson, Giovanni Motta, Lillian Zhou, Franoise Beaufays, Arnaud Benard, Dhruv Guliani, Andreas Kabel, Nikhil Khare, Tamar Lucassen, Petr Zadrazil, Harry Zhang |
| 2019 | ICASSP | Streaming End-to-end Speech Recognition for Mobile Devices. | Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, Qiao Liang, Deepti Bhatia, Yuan Shangguan, Bo Li, Golan Pundak, Khe Chai Sim, Tom Bagby, Shuo-Yiin Chang, Kanishka Rao, Alexander Gruenstein |
| 2019 | ICASSP | Improving CTC Using Stimulated Learning for Sequence Modeling. | Jahn Heymann, Khe Chai Sim, Bo Li |
| 2019 | Interspeech | An Investigation into On-Device Personalization of End-to-End Automatic Speech Recognition Models. | Khe Chai Sim, Petr Zadrazil, Franoise Beaufays |
| 2018 | ICASSP | Understanding Recurrent Neural State Using Memory Signatures. | Skanda Koppula, Khe Chai Sim, Kean K. Chin |
| 2018 | ICASSP | Multi-Dialect Speech Recognition with a Single Sequence-to-Sequence Model. | Bo Li, Tara N. Sainath, Khe Chai Sim, Michiel Bacchiani, Eugene Weinstein, Patrick Nguyen, Zhifeng Chen, Yanghui Wu, Kanishka Rao |
| 2018 | ICASSP | learning Effective Factorized Hidden Layer Bases Using Student-Teacher Training for LSTM Acoustic Model Adaptation. | Lahiru Samarakoon, Brian Mak, Khe Chai Sim |
| 2018 | Interspeech | Domain Adaptation Using Factorized Hidden Layer for Robust Automatic Speech Recognition. | Khe Chai Sim, Arun Narayanan, Ananya Misra, Anshuman Tripathi, Golan Pundak, Tara N. Sainath, Parisa Haghani, Bo Li, Michiel Bacchiani |
| 2017 | ASRU | Improving the efficiency of forward-backward algorithm using batched computation in TensorFlow. | Khe Chai Sim, Arun Narayanan, Tom Bagby, Tara N. Sainath, Michiel Bacchiani |
| 2017 | ICASSP | An investigation into learning effective speaker subspaces for robust unsupervised DNN adaptation. | Lahiru Samarakoon, Khe Chai Sim, Brian Mak |
| 2017 | Interspeech | Acoustic Modeling for Google Home. | Bo Li, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim, Olivier Siohan, Mitchel Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
| 2017 | Interspeech | Learning Factorized Transforms for Unsupervised Adaptation of LSTM-RNN Acoustic Models. | Lahiru Samarakoon, Brian Mak, Khe Chai Sim |
| 2017 | Interspeech | An Efficient Phone N-Gram Forward-Backward Computation Using Dense Matrix Multiplication. | Khe Chai Sim, Arun Narayanan |
| 2016 | ICASSP | Joint acoustic factor learning for robust deep neural network based automatic speech recognition. | Souvik Kundu, Gautam Mantena, Yanmin Qian, Tian Tan, Marc Delcroix, Khe Chai Sim |
| 2016 | ICASSP | On combining i-vectors and discriminative adaptation methods for unsupervised speaker normalization in DNN acoustic models. | Lahiru Samarakoon, Khe Chai Sim |
| 2016 | ICASSP | Speaker-aware training of LSTM-RNNS for acoustic modelling. | Tian Tan, Yanmin Qian, Dong Yu, Souvik Kundu, Liang Lu, Khe Chai Sim, Xiong Xiao, Yu Zhang |
| 2016 | ICASSP | Towards implicit complexity control using variable-depth deep neural networks for automatic speech recognition. | Shawn Tan, Khe Chai Sim |
| 2016 | Interspeech | Incorporating a Generative Front-End Layer to Deep Neural Network for Noise Robust Automatic Speech Recognition. | Souvik Kundu, Khe Chai Sim, Mark J. F. Gales |
| 2016 | Interspeech | Microphone Distance Adaptation Using Cluster Adaptive Training for Robust Far Field Speech Recognition. | Animesh Prasad, Khe Chai Sim |
| 2016 | Interspeech | Subspace LHUC for Fast Adaptation of Deep Neural Network Acoustic Models. | Lahiru Samarakoon, Khe Chai Sim |
| 2016 | Interspeech | Multi-Attribute Factorized Hidden Layer Adaptation for DNN Acoustic Models. | Lahiru Samarakoon, Khe Chai Sim |
| 2016 | Interspeech | Stimulated Deep Neural Network for Speech Recognition. | Chunyang Wu, Penny Karanasou, Mark J. F. Gales, Khe Chai Sim |
| 2015 | ASRU | Learning factorized feature transforms for speaker normalization. | Lahiru Samarakoon, Khe Chai Sim |
| 2015 | ASRU | On constructing and analysing an interpretable brain model for the DNN based on hidden activity patterns. | Khe Chai Sim |
| 2015 | ASRU | Improving the interpretability of deep neural networks with stimulated learning. | Shawn Tan, Khe Chai Sim, Mark J. F. Gales |
| 2015 | ICASSP | An investigation of augmenting speaker representations to improve speaker normalisation for DNN-based speech recognition. | Hengguan Huang, Khe Chai Sim |
| 2014 | COLING | A Beam-Search Decoder for Disfluency Detection. | Xuancong Wang, Hwee Tou Ng, Khe Chai Sim |
| 2014 | EMNLP | Combining Punctuation and Disfluency Prediction: An Empirical Study. | Xuancong Wang, Khe Chai Sim, Hwee Tou Ng |
| 2014 | ICASSP | Second order vector taylor series based robust speech recognition. | Suliang Bu, Yanmin Qian, Khe Chai Sim, Yongbin You, Kai Yu |
| 2014 | ICASSP | An ideal hidden-activation mask for deep neural networks based noise-robust speech recognition. | Bo Li, Khe Chai Sim |
| 2014 | ICASSP | On combining DNN and GMM with unsupervised speaker adaptation for robust automatic speech recognition. | Shilin Liu, Khe Chai Sim |
| 2014 | ICASSP | Refinements of regression-based context-dependent modelling of deep neural networks for automatic speech recognition. | Guangsen Wang, Khe Chai Sim |
| 2014 | Interspeech | Modeling long temporal contexts for robust DNN-based speech recognition. | Bo Li, Khe Chai Sim |
| 2014 | Interspeech | Joint adaptation and adaptive training of TVWR for robust automatic speech recognition. | Shilin Liu, Khe Chai Sim |
| 2013 | ASRU | Improving robustness of deep neural networks via spectral masking for automatic speech recognition. | Bo Li, Khe Chai Sim |
| 2013 | ASRU | Multi-stream temporally varying weight regression for cross-lingual speech recognition. | Shilin Liu, Khe Chai Sim |
| 2013 | ASRU | Context-dependent modelling of deep neural network using logistic regression. | Guangsen Wang, Khe Chai Sim |
| 2013 | ICASSP | Noise adaptive front-end normalization based on Vector Taylor Series for Deep Neural Networks in robust speech recognition. | Bo Li, Khe Chai Sim |
| 2013 | ICASSP | Approximated Parallel Model Combination for efficient noise-robust speech recognition. | Khe Chai Sim |
| 2013 | Interspeech | An investigation of spectral restoration algorithms for deep neural networks based noise robust speech recognition. | Bo Li, Yu Tsao, Khe Chai Sim |
| 2013 | Interspeech | Parameter clustering for temporally varying weight regression for automatic speech recognition. | Shilin Liu, Khe Chai Sim |
| 2013 | Interspeech | An investigation of temporally varying weight regression for noise robust speech recognition. | Shilin Liu, Khe Chai Sim |
| 2013 | Interspeech | Integrating conditional random fields and joint multi-gram model with syllabic features for grapheme-to-phone conversion. | Xiaoxuan Wang, Khe Chai Sim |
| 2012 | ACL | Probabilistic Integration of Partial Lexical Information for Noise Robust Haptic Voice Recognition. | Khe Chai Sim |
| 2012 | ICASSP | Implicit trajectory modelling using temporally varying weight regression for automatic speech recognition. | Shilin Liu, Khe Chai Sim |
| 2012 | ICASSP | An investigation of tied-mixture GMM based triphone state clustering. | Guangsen Wang, Khe Chai Sim |
| 2012 | ICMI | Design and implementation of the note-taking style haptic voice recognition for mobile devices. | Seungwhan Moon, Khe Chai Sim |
| 2012 | ICMI | Speak-as-you-swipe (SAYS): a multimodal interface combining speech and gesture keyboard synchronously for continuous mobile text entry. | Khe Chai Sim |
| 2012 | ICMI | ICMI'12 grand challenge: haptic voice recognition. | Khe Chai Sim, Shengdong Zhao, Kai Yu, Hank Liao |
| 2012 | ICMI | Improving mandarin predictive text input by augmenting pinyin initials with speech and tonal information. | Guangsen Wang, Bo Li, Shilin Liu, Xuancong Wang, Xiaoxuan Wang, Khe Chai Sim |
| 2012 | Interspeech | A Weighted Combination of Speech with Text-based Models for Arabic Diacritization. | Aisha S. Azim, Xiaoxuan Wang, Khe Chai Sim |
| 2012 | Interspeech | A Two-stage Speaker Adaptation Approach for Subspace Gaussian Mixture Model based Nonnative Speech Recognition. | Bo Li, Khe Chai Sim |
| 2012 | Interspeech | Dynamic Conditional Random Fields for Joint Sentence Boundary and Punctuation Prediction. | Xuancong Wang, Hwee Tou Ng, Khe Chai Sim |
| 2011 | ASRU | A Trajectory-based Parallel Model Combination with a unified static and dynamic parameter compensation for noisy speech recognition. | Khe Chai Sim, Minh-Thang Luong |
| 2011 | Interspeech | Sequential Classification Criteria for NNs in Automatic Speech Recognition. | Guangsen Wang, Khe Chai Sim |
| 2011 | Interspeech | Comparison of Smoothing Techniques for Robust Context Dependent Acoustic Modelling in Hybrid NN/HMM Systems. | Guangsen Wang, Khe Chai Sim |
| 2010 | ICASSP | A minimum variance asynchronous Detection Error Trade-off performance analysis for multi-class detection problems. | Khe Chai Sim |
| 2010 | ICASSP | Adaptive score fusion using Weighted Logistic Linear Regression for spoken language recognition. | Khe Chai Sim, Kong-Aik Lee |
| 2010 | Interspeech | Comparison of discriminative input and output transformations for speaker adaptation in the hybrid NN/HMM systems. | Bo Li, Khe Chai Sim |
| 2010 | Interspeech | Hidden logistic linear regression for support vector machine based phone verification. | Bo Li, Khe Chai Sim |
| 2010 | Interspeech | Probabilistic state clustering using conditional random field for context-dependent acoustic modelling. | Khe Chai Sim |
| 2010 | Interspeech | Semi-parametric trajectory modelling using temporally varying feature mapping for speech recognition. | Khe Chai Sim, Shilin Liu |
| 2009 | ASRU | Discriminative Product-of-Expert acoustic mapping for cross-lingual phone recognition. | Khe Chai Sim |
| 2009 | ICASSP | The I4U system in NIST 2008 speaker recognition evaluation. | Haizhou Li, Bin Ma, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Changhuai You, Rong Tong, Ismo Krkkinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li, Li-Rong Dai, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin |
| 2009 | Interspeech | Stream-based context-sensitive phone mapping for cross-lingual speech recognition. | Khe Chai Sim, Haizhou Li |
| 2008 | ICASSP | Robust phone set mapping using decision tree clustering for cross-lingual phone recognition. | Khe Chai Sim, Haizhou Li |
| 2008 | Interspeech | Context-sensitive probabilistic phone mapping model for cross-lingual speech recognition. | Khe Chai Sim, Haizhou Li |
| 2008 | PACLIC | NIST 2007 Language Recognition Evaluation: From the Perspective of IIR. | Haizhou Li, Bin Ma, Kong-Aik Lee, Khe Chai Sim, Hanwu Sun, Rong Tong, Donglai Zhu, Changhuai You |
| 2008 | SIGIR | A lattice-based approach to query-by-example spoken document retrieval. | Tee Kiah Chia, Khe Chai Sim, Haizhou Li, Hwee Tou Ng |
| 2007 | ACL | Semantic Transliteration of Personal Names. | Haizhou Li, Khe Chai Sim, Jin-Shea Kuo, Minghui Dong |
| 2007 | ICASSP | Consensus Network Decoding for Statistical Machine Translation System Combination. | Khe Chai Sim, William J. Byrne, Mark J. F. Gales, Hichem Sahbi, Philip C. Woodland |
| 2007 | ICASSP | Improving Speech Transcription for Mandarin-English Translation. | Marcus Tomalin, Mark J. F. Gales, X. Andrew Liu, Khe Chai Sim, Rohit Sinha, Lan Wang, Philip C. Woodland, Kai Yu |
| 2007 | Interspeech | Fusion of contrastive acoustic models for parallel phonotactic spoken language identification. | Khe Chai Sim, Haizhou Li |
| 2006 | ICASSP | The Cu-Htk Mandarin Broadcast News Transcription System. | Rohit Sinha, Mark J. F. Gales, Do Yeong Kim, X. Andrew Liu, Khe Chai Sim, Philip C. Woodland |
| 2005 | ICASSP | Development of the CUHTK 2004 Mandarin Conversational Telephone Speech Transcription System. | Mark J. F. Gales, Bin Jia, X. Andrew Liu, Khe Chai Sim, Philip C. Woodland, Kai Yu |
| 2005 | ICASSP | Development of the CU-HTK 2004 Broadcast News Transcription Systems. | Do Yeong Kim, Ho Yin Chan, Gunnar Evermann, Mark J. F. Gales, David Mrva, Khe Chai Sim, Philip C. Woodland |
| 2005 | ICASSP | Investigation of Acoustic Modeling Techniques for LVCSR Systems. | Xunying Liu, Mark J. F. Gales, Khe Chai Sim, Kai Yu |
| 2005 | ICASSP | Adaptation of Precision Matrix Models on Large Vocabulary Continuous Speech Recognition. | Khe Chai Sim, Mark J. F. Gales |
| 2005 | Interspeech | Temporally varying model parameters for large vocabulary continuous speech recognition. | Khe Chai Sim, Mark J. F. Gales |
| 2004 | ICASSP | Basis superposition precision matrix modelling for large vocabulary continuous speech recognition. | Khe Chai Sim, Mark J. F. Gales |