| 2025 | ICASSP | AMuSE: Attentive Multilingual Speech Encoding for Zero-Prior ASR. | Ashutosh Varshney, Debmalya Chakrabarty, Akshat Jaiswal, Harish Arsikere, Abhinav Jain, Swayambhu Nath Ray, Frederick Weber, Anand Mohan, Prantik Sen, Garima Lalwani, Sambuddha Bhattacharya, Sri Garimella |
| 2025 | Interspeech | DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation. | Prabash Reddy Male, Swayambhu Nath Ray, Harish Arsikere, Akshat Jaiswal, Prakhar Swarup, Prantik Sen, Debmalya Chakrabarty, K. V. Vijay Girish, Nikhil Bhave, Frederick Weber, Sambuddha Bhattacharya, Sri Garimella |
| 2021 | ICASSP | REDAT: Accent-Invariant Representation for End-To-End ASR by Domain Adversarial Training with Relabeling. | Hu Hu, Xuesong Yang, Zeynab Raeesy, Jinxi Guo, Gokce Keskin, Harish Arsikere, Ariya Rastrow, Andreas Stolcke, Roland Maas |
| 2021 | ICASSP | Joint ASR and Language Identification Using RNN-T: An Efficient Approach to Dynamic Language Switching. | Surabhi Punjabi, Harish Arsikere, Zeynab Raeesy, Chander Chandak, Nikhil Bhave, Ankish Bansal, Markus Mller, Sergio Murillo, Ariya Rastrow, Andreas Stolcke, Jasha Droppo, Sri Garimella, Roland Maas, Mat Hans, Athanasios Mouchtaris, Siegfried Kunzmann |
| 2021 | Interspeech | Listen with Intent: Improving Speech Recognition with Audio-to-Intent Front-End. | Swayambhu Nath Ray, Minhua Wu, Anirudh Raju, Pegah Ghahremani, Raghavendra Bilgi, Milind Rao, Harish Arsikere, Ariya Rastrow, Andreas Stolcke, Jasha Droppo |
| 2020 | Interspeech | Improved Training Strategies for End-to-End Speech Recognition in Digital Voice Assistants. | Hitesh Tulsiani, Ashtosh Sapru, Harish Arsikere, Surabhi Punjabi, Sri Garimella |
| 2019 | ASRU | Language Model Bootstrapping Using Neural Machine Translation for Conversational Speech Recognition. | Surabhi Punjabi, Harish Arsikere, Sri Garimella |
| 2019 | Interspeech | Multi-Dialect Acoustic Modeling Using Phone Mapping and Online i-Vectors. | Harish Arsikere, Ashtosh Sapru, Sri Garimella |
| 2017 | Interspeech | Robust Online i-Vectors for Unsupervised Adaptation of DNN Acoustic Models: A Study in the Context of Digital Voice Assistants. | Harish Arsikere, Sri Garimella |
| 2016 | ICASSP | Novel acoustic features for automatic dialog-act tagging. | Harish Arsikere, Arunasish Sen, A. P. Prathosh, Vivek Tyagi |
| 2016 | Interspeech | Speaker Verification Using Short Utterances with DNN-Based Estimation of Subglottal Acoustic Features. | Jinxi Guo, Gary Yeung, Deepak Muralidharan, Harish Arsikere, Amber Afshan, Abeer Alwan |
| 2015 | Interspeech | Stylex: a corpus of educational videos for research on speaking styles and their impact on engagement and learning. | Harish Arsikere, Sonal Patil, Ranjeet Kumar, Kundan Shrivastava, Om Deshmukh |
| 2015 | Interspeech | Age-dependent height estimation and speaker normalization for children's speech using the first three subglottal resonances. | Jinxi Guo, Rohit Paturi, Gary Yeung, Steven M. Lulich, Harish Arsikere, Abeer Alwan |
| 2015 | Interspeech | Acoustic stress detection for improved navigation of educational videos. | Sonal Patil, Harish Arsikere, Om Deshmukh |
| 2015 | IUI | Content-driven Multi-modal Techniques for Non-linear Video Navigation. | Kuldeep Yadav, Kundan Shrivastava, S. Mohana Prasad, Harish Arsikere, Sonal Patil, Ranjeet Kumar, Om Deshmukh |
| 2014 | ICASSP | Frequency warping using subglottal resonances: Complementarity with VTLN and robustness to additive noise. | Harish Arsikere, Abeer Alwan |
| 2014 | ICASSP | Computationally-efficient endpointing features for natural spoken interaction with personal-assistant systems. | Harish Arsikere, Elizabeth Shriberg, Umut Ozertem |
| 2014 | Interspeech | Speaker recognition via fusion of subglottal features and MFCCs. | Harish Arsikere, Hitesh Anand Gupta, Abeer Alwan |
| 2014 | Interspeech | The relationship between the second subglottal resonance and vowel class, standing height, trunk length, and F0 variation for Mandarin speakers. | Jinxi Guo, Angli Liu, Harish Arsikere, Abeer Alwan, Steven M. Lulich |
| 2013 | ICASSP | Non-linear frequency warping for VTLN using subglottal resonances and the third formant frequency. | Harish Arsikere, Steven M. Lulich, Abeer Alwan |
| 2012 | ICASSP | Automatic height estimation using the second subglottal resonance. | Harish Arsikere, Gary K. F. Leung, Steven M. Lulich, Abeer Alwan |
| 2012 | Interspeech | Automatic estimation of the first two subglottal resonances in children's speech with application to speaker normalization in limited-data conditions. | Harish Arsikere, Gary K. F. Leung, Steven M. Lulich, Abeer Alwan |
| 2011 | ICASSP | Automatic estimation of the second subglottal resonance from natural speech. | Harish Arsikere, Steven M. Lulich, Abeer Alwan |
| 2011 | Interspeech | Analysis and Automatic Estimation of Children's Subglottal Resonances. | Steven M. Lulich, Harish Arsikere, John R. Morton, Gary K. F. Leung, Abeer Alwan, Mitchell Sommers |