| 2025 | ICASSP | Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition. | Takaaki Hori, Martin Kocour, Adnan Haider, Erik McDermott, Xiaodan Zhuang |
| 2023 | ICASSP | Variable Attention Masking for Configurable Transformer Transducer Speech Recognition. | Pawel Swietojanski, Stefan Braun, Dogan Can, Thiago Fraga da Silva, Arnab Ghoshal, Takaaki Hori, Roger Hsiao, Henry Mason, Erik McDermott, Honza Silovsky, Ruchir Travadi, Xiaodan Zhuang |
| 2023 | Interspeech | Approximate Nearest Neighbour Phrase Mining for Contextual Speech Recognition. | Maurits J. R. Bleeker, Pawel Swietojanski, Stefan Braun, Xiaodan Zhuang |
| 2020 | ICASSP | SNDCNN: Self-Normalizing Deep CNNs with Scaled Exponential Linear Units for Speech Recognition. | Zhen Huang, Tim Ng, Leo Liu, Henry Mason, Xiaodan Zhuang, Daben Liu |
| 2019 | ICASSP | Exploring Retraining-free Speech Recognition for Intra-sentential Code-switching. | Zhen Huang, Xiaodan Zhuang, Daben Liu, Xiaoqiang Xiao, Yuchen Zhang, Sabato Marco Siniscalchi |
| 2017 | Interspeech | Improving DNN Bluetooth Narrowband Acoustic Models by Cross-Bandwidth and Cross-Lingual Initialization. | Xiaodan Zhuang, Arnab Ghoshal, Antti-Veikko Rosti, Matthias Paulik, Daben Liu |
| 2014 | CVPR | Zero-Shot Event Detection Using Multi-modal Fusion of Weakly Supervised Concepts. | Shuang Wu, Sravanthi Bondugula, Florian Luisier, Xiaodan Zhuang, Pradeep Natarajan |
| 2014 | DAS | Text Classification via iVector Based Feature Representation. | Shengxin Zha, Xujun Peng, Huaigu Cao, Xiaodan Zhuang, Pradeep Natarajan, Prem Natarajan |
| 2014 | ICASSP | Text detection and recognition in natural scenes and consumer videos. | Arpit Jain, Xujun Peng, Xiaodan Zhuang, Pradeep Natarajan, Huaigu Cao |
| 2014 | ICASSP | Effective representations for leveraging language content in multimedia event detection. | Shuang Wu, Xiaodan Zhuang, Pradeep Natarajan |
| 2013 | Interspeech | Audio self organized units for high-level event detection. | Xiaodan Zhuang, Shuang Wu, Pradeep Natarajan, Rohit Prasad, Prem Natarajan |
| 2013 | Interspeech | Probabilistic trainable segmenter for call center audio using multiple features. | Nina Zinovieva, Xiaodan Zhuang, Pat Peterson, Joe Alwan, Rohit Prasad |
| 2013 | WACV | Scene image categorization and video event detection using Naive Bayes Nearest Neighbor. | Shiv Naga Prasad Vitaladevuni, Pradeep Natarajan, Shuang Wu, Xiaodan Zhuang, Rohit Prasad, Premkumar Natarajan |
| 2012 | CVPR | Multimodal feature fusion for robust event detection in web videos. | Pradeep Natarajan, Shuang Wu, Shiv Naga Prasad Vitaladevuni, Xiaodan Zhuang, Stavros Tsakalidis, Unsang Park, Rohit Prasad, Premkumar Natarajan |
| 2012 | ECCV | Multi-channel Shape-Flow Kernel Descriptors for Robust Video Event Detection and Retrieval. | Pradeep Natarajan, Shuang Wu, Shiv Naga Prasad Vitaladevuni, Xiaodan Zhuang, Unsang Park, Rohit Prasad, Premkumar Natarajan |
| 2012 | ICASSP | Improving faster-than-real-time human acoustic event detection by saliency-maximized audio visualization. | Kai-Hsiang Lin, Xiaodan Zhuang, Camille Goudeseune, Sarah King, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2012 | Interspeech | Robust Event Detection From Spoken Content In Consumer Domain Videos. | Stavros Tsakalidis, Xiaodan Zhuang, Roger Hsiao, Shuang Wu, Pradeep Natarajan, Rohit Prasad, Prem Natarajan |
| 2012 | Interspeech | Compact Audio Representation for Event Detection in Consumer Media. | Xiaodan Zhuang, Stavros Tsakalidis, Shuang Wu, Pradeep Natarajan, Rohit Prasad, Prem Natarajan |
| 2011 | ICASSP | Improving acoustic event detection using generalizable visual features and multi-modality modeling. | Po-Sen Huang, Xiaodan Zhuang, Mark Hasegawa-Johnson |
| 2011 | ICASSP | Synthesizing visual speech trajectory with minimum generation error. | Lijuan Wang, Yi-Jian Wu, Xiaodan Zhuang, Frank K. Soong |
| 2010 | Interspeech | FSM-based pronunciation modeling using articulatory phonological code. | Chi Hu, Xiaodan Zhuang, Mark Hasegawa-Johnson |
| 2010 | Interspeech | A minimum converted trajectory error (MCTE) approach to high quality speech-to-lips conversion. | Xiaodan Zhuang, Lijuan Wang, Frank K. Soong, Mark Hasegawa-Johnson |
| 2009 | ICASSP | Long-time span acoustic activity analysis from far-field sensors in smart homes. | Jing Huang, Xiaodan Zhuang, Vit Libal, Gerasimos Potamianos |
| 2009 | ICASSP | Acoustic fall detection using Gaussian mixture models and GMM supervectors. | Xiaodan Zhuang, Jing Huang, Gerasimos Potamianos, Mark Hasegawa-Johnson |
| 2009 | Interspeech | Articulatory phonological code for word classification. | Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman |
| 2008 | ICASSP | Feature analysis and selection for acoustic event detection. | Xiaodan Zhuang, Xi Zhou, Thomas S. Huang, Mark Hasegawa-Johnson |
| 2008 | ICPR | A novel Gaussianized vector representation for natural scene categorization. | Xi Zhou, Xiaodan Zhuang, Hao Tang, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2008 | ICPR | Face age estimation using patch-based hidden Markov model supervectors. | Xiaodan Zhuang, Xi Zhou, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2008 | Interspeech | The entropy of the articulatory phonological code: recognizing gestures from tract variables. | Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis M. Goldstein, Elliot Saltzman |