| 2024 | VTC | 3D YOLO-SM: End-to-End Approach for Real-time Traffic Light Detection and Recognition in Complex Scenarios. | Kshitiz Kumar, D. Santhosh Reddy, P. Rajalakshmi |
| 2022 | ICASSP | Maximizing Audio Event Detection Model Performance on Small Datasets Through Knowledge Transfer, Data Augmentation, and Pretraining: an Ablation Study. | Daniel Tompkins, Kshitiz Kumar, Jian Wu |
| 2021 | ICASSP | Multi-Dialect Speech Recognition in English Using Attention on Ensemble of Experts. | Amit Das, Kshitiz Kumar, Jian Wu |
| 2021 | Interspeech | Sequence-Level Confidence Classifier for ASR Utterance Accuracy and Application to Acoustic Models. | Amber Afshan, Kshitiz Kumar, Jian Wu |
| 2020 | Interspeech | Transfer Learning Approaches for Streaming End-to-End Speech Recognition System. | Vikas Joshi, Rui Zhao, Rupesh R. Mehta, Kshitiz Kumar, Jinyu Li |
| 2020 | Interspeech | 1-D Row-Convolution LSTM: Fast Streaming ASR at Accuracy Parity with LC-BLSTM. | Kshitiz Kumar, Chaojun Liu, Yifan Gong, Jian Wu |
| 2020 | Interspeech | Bandpass Noise Generation and Augmentation for Unified ASR. | Kshitiz Kumar, Bo Ren, Yifan Gong, Jian Wu |
| 2020 | Interspeech | Fast and Slow Acoustic Model. | Kshitiz Kumar, Emilian Stoimenov, Hosam Khalil, Jian Wu |
| 2019 | ICASSP | Word Characters and Phone Pronunciation Embedding for ASR Confidence Classifier. | Kshitiz Kumar, Tasos Anastasakos, Yifan Gong |
| 2019 | ICASSP | Static and Dynamic State Predictions for Acoustic Model Combination. | Kshitiz Kumar, Yifan Gong |
| 2018 | ISNN | A Comparative Study of Spatial Speech Separation Techniques to Improve Speech Recognition. | Xinhui Zhou, Chiman Kwan, Bulent Ayhan, Chanwoo Kim, Kshitiz Kumar, Richard M. Stern |
| 2017 | ICASSP | Extended low-rank plus diagonal adaptation for deep and recurrent neural networks. | Yong Zhao, Jinyu Li, Kshitiz Kumar, Yifan Gong |
| 2016 | ICASSP | Non-negative intermediate-layer DNN adaptation for a 10-KB speaker adaptation profile. | Kshitiz Kumar, Chaojun Liu, Yifan Gong |
| 2016 | ICASSP | Investigations on speaker adaptation of LSTM RNN models for speech recognition. | Chaojun Liu, Yongqiang Wang, Kshitiz Kumar, Yifan Gong |
| 2015 | Interspeech | Confidence-features and confidence-scores for ASR applications in arbitration and DNN speaker adaptation. | Kshitiz Kumar, Ziad Al Bawab, Yong Zhao, Chaojun Liu, Benot Dumoulin, Yifan Gong |
| 2015 | Interspeech | Delta-melspectra features for noise robustness to DNN-based ASR systems. | Kshitiz Kumar, Chaojun Liu, Yifan Gong |
| 2015 | Interspeech | Intermediate-layer DNN adaptation for offline and session-based iterative speaker adaptation. | Kshitiz Kumar, Chaojun Liu, Kaisheng Yao, Yifan Gong |
| 2014 | Interspeech | Normalization of ASR confidence classifier scores via confidence mapping. | Kshitiz Kumar, Chaojun Liu, Yifan Gong |
| 2013 | ICASSP | Predicting speech recognition confidence using deep learning with word identity and score features. | Po-Sen Huang, Kshitiz Kumar, Chaojun Liu, Yifan Gong, Li Deng |
| 2011 | ICASSP | Binaural sound source separation motivated by auditory processing. | Chanwoo Kim, Kshitiz Kumar, Richard M. Stern |
| 2011 | ICASSP | Delta-spectral cepstral coefficients for robust speech recognition. | Kshitiz Kumar, Chanwoo Kim, Richard M. Stern |
| 2011 | ICASSP | An iterative least-squares technique for dereverberation. | Kshitiz Kumar, Bhiksha Raj, Rita Singh, Richard M. Stern |
| 2011 | ICASSP | Gammatone sub-band magnitude-domain dereverberation for ASR. | Kshitiz Kumar, Rita Singh, Bhiksha Raj, Richard M. Stern |
| 2010 | ICASSP | Maximum-likelihood-based cepstral inverse filtering for blind speech dereverberation. | Kshitiz Kumar, Richard M. Stern |
| 2009 | ASRU | Robust speech recognition using a Small Power Boosting algorithm. | Chanwoo Kim, Kshitiz Kumar, Richard M. Stern |
| 2009 | CVPR | Audio-visual speech synchronization detection using a bimodal linear prediction model. | Kshitiz Kumar, Jir Navrtil, Etienne Marcheret, Vit Libal, Ganesh N. Ramaswamy, Gerasimos Potamianos |
| 2009 | Interspeech | Signal separation for robust speech recognition based on phase difference information obtained in the frequency domain. | Chanwoo Kim, Kshitiz Kumar, Bhiksha Raj, Richard M. Stern |
| 2009 | Interspeech | Robust audio-visual speech synchrony detection by generalized bimodal linear prediction. | Kshitiz Kumar, Jir Navrtil, Etienne Marcheret, Vit Libal, Gerasimos Potamianos |
| 2008 | ICASSP | Environment-invariant compensation for reverberation using linear post-filtering for minimum distortion. | Kshitiz Kumar, Richard M. Stern |
| 2007 | ICASSP | Profile View Lip Reading. | Kshitiz Kumar, Tsuhan Chen, Richard M. Stern |