| 2025 | ICASSP | Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition. | Keqi Deng, Jinxi Guo, Yingyi Ma, Niko Moritz, Philip C. Woodland, Ozlem Kalinli, Mike Seltzer |
| 2025 | ICASSP | A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data. | Minh Tran, Yutong Pang, Debjyoti Paul, Laxmi Pandey, Kevin Jiang, Jinxi Guo, Ke Li, Shun Zhang, Xuedong Zhang, Xin Lei |
| 2025 | Interspeech | R2S: Real-to-Synthetic Representation Learning for Training Speech Recognition Models on Synthetic Data. | Minh Tran, Debjyoti Paul, Yutong Pang, Laxmi Pandey, Jinxi Guo, Ke Li, Shun Zhang, Xuedong Zhang, Xin Lei |
| 2024 | ICASSP | Prompting Large Language Models with Speech Recognition Abilities. | Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, Christian Fuegen, Mike Seltzer |
| 2024 | ICASSP | Effective Internal Language Model Training and Fusion for Factorized Transducer Model. | Jinxi Guo, Niko Moritz, Yingyi Ma, Frank Seide, Chunyang Wu, Jay Mahadeokar, Ozlem Kalinli, Christian Fuegen, Mike Seltzer |
| 2024 | ICASSP | Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of a Multilingual ASR Model. | Jiamin Xie, Ke Li, Jinxi Guo, Andros Tjandra, Yuan Shangguan, Leda Sari, Chunyang Wu, Junteng Jia, Jay Mahadeokar, Ozlem Kalinli |
| 2023 | ICASSP | Improving fast-slow Encoder based Transducer with Streaming Deliberation. | Ke Li, Jay Mahadeokar, Jinxi Guo, Yangyang Shi, Gil Keren, Ozlem Kalinli, Michael L. Seltzer, Duc Le |
| 2023 | Interspeech | Biased Self-supervised Learning for ASR. | Florian L. Kreyssig, Yangyang Shi, Jinxi Guo, Leda Sari, Abdel-rahman Mohamed, Philip C. Woodland |
| 2022 | ICASSP | VADOI: Voice-Activity-Detection Overlapping Inference for End-To-End Long-Form Speech Recognition. | Jinhan Wang, Xiaosu Tong, Jinxi Guo, Di He, Roland Maas |
| 2021 | ICASSP | REDAT: Accent-Invariant Representation for End-To-End ASR by Domain Adversarial Training with Relabeling. | Hu Hu, Xuesong Yang, Zeynab Raeesy, Jinxi Guo, Gokce Keskin, Harish Arsikere, Ariya Rastrow, Andreas Stolcke, Roland Maas |
| 2020 | Interspeech | Variable Frame Rate-Based Data Augmentation to Handle Speaking-Style Variability for Automatic Speaker Verification. | Amber Afshan, Jinxi Guo, Soo Jin Park, Vijay Ravi, Alan McCree, Abeer Alwan |
| 2020 | Interspeech | Efficient Minimum Word Error Rate Training of RNN-Transducer for End-to-End Speech Recognition. | Jinxi Guo, Gautam Tiwari, Jasha Droppo, Maarten Van Segbroeck, Che-Wei Huang, Andreas Stolcke, Roland Maas |
| 2019 | ICASSP | A Spelling Correction Model for End-to-end Speech Recognition. | Jinxi Guo, Tara N. Sainath, Ron J. Weiss |
| 2018 | ICASSP | Time-Delayed Bottleneck Highway Networks Using a DFT Feature for Keyword Spotting. | Jinxi Guo, Ken'ichi Kumatani, Ming Sun, Minhua Wu, Anirudh Raju, Nikko Strom, Arindam Mandal |
| 2018 | Interspeech | Effectiveness of Voice Quality Features in Detecting Depression. | Amber Afshan, Jinxi Guo, Soo Jin Park, Vijay Ravi, Jonathan Flint, Abeer Alwan |
| 2018 | Interspeech | Filter Sampling and Combination CNN (FSC-CNN): A Compact CNN Model for Small-footprint ASR Acoustic Modeling Using Raw Waveforms. | Jinxi Guo, Ning Xu, Xin Chen, Yang Shi, Kaiyuan Xu, Abeer Alwan |
| 2017 | Interspeech | CNN-Based Joint Mapping of Short and Long Utterance i-Vectors for Speaker Verification Using Short Utterances. | Jinxi Guo, Usha Amrutha Nookala, Abeer Alwan |
| 2017 | Interspeech | Attention Based CLDNNs for Short-Duration Acoustic Scene Classification. | Jinxi Guo, Ning Xu, Li-Jia Li, Abeer Alwan |
| 2016 | Interspeech | Speaker Verification Using Short Utterances with DNN-Based Estimation of Subglottal Acoustic Features. | Jinxi Guo, Gary Yeung, Deepak Muralidharan, Harish Arsikere, Amber Afshan, Abeer Alwan |
| 2016 | Interspeech | Speaker Identity and Voice Quality: Modeling Human Responses and Automatic Speaker Recognition. | Soo Jin Park, Caroline Sigouin, Jody Kreiman, Patricia A. Keating, Jinxi Guo, Gary Yeung, Fang-Yu Kuo, Abeer Alwan |
| 2015 | Interspeech | Age-dependent height estimation and speaker normalization for children's speech using the first three subglottal resonances. | Jinxi Guo, Rohit Paturi, Gary Yeung, Steven M. Lulich, Harish Arsikere, Abeer Alwan |
| 2014 | Interspeech | The relationship between the second subglottal resonance and vowel class, standing height, trunk length, and F0 variation for Mandarin speakers. | Jinxi Guo, Angli Liu, Harish Arsikere, Abeer Alwan, Steven M. Lulich |