| 2025 | ASRU | MADASR 2.0: Multi-Lingual Multi-Dialect ASR Challenge in 8 Indian Languages. | Saurabh Kumar, Sumit Sharma, Deekshitha G, Abhayjeet Singh, Amartyaveer, Sathvik Udupa, Sandhya Badiger, Sanjeev Khudanpur, Sunayana Sitaram, Srinivasan Umesh, Bhuvana Ramabhadran, Brian Kingsbury, Hema A. Murthy, Srikanth S. Narayanan, Howard Lakougna, Prasanta Kumar Ghosh |
| 2025 | EMNLP | EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion. | Advait Joglekar, Divyanshu Singh, Rooshil Rohit Bhatia, Srinivasan Umesh |
| 2024 | ICASSP | Stable Distillation: Regularizing Continued Pre-Training for Low-Resource Automatic Speech Recognition. | Ashish Seth, Sreyan Ghosh, Srinivasan Umesh, Dinesh Manocha |
| 2024 | ICASSP | FusDom: Combining in-Domain and Out-of-Domain Knowledge for Continuous Self-Supervised Learning. | Ashish Seth, Sreyan Ghosh, Srinivasan Umesh, Dinesh Manocha |
| 2024 | Interspeech | All Ears: Building Self-Supervised Learning based ASR models for Indian Languages at scale. | Vasista Sai Lodagala, Abhishek Biswas, Shoutrik Das, Jordan Fernandes, Srinivasan Umesh |
| 2023 | ASRU | Towards Developing State-of-The-Art TTS Synthesisers for 13 Indian Languages with Signal Processing Aided Alignments. | Anusha Prakash, Srinivasan Umesh, Hema A. Murthy |
| 2023 | ICASSP | MAST: Multiscale Audio Spectrogram Transformers. | Sreyan Ghosh, Ashish Seth, Srinivasan Umesh, Dinesh Manocha |
| 2023 | ICASSP | Data2vec-Aqc: Search for the Right Teaching Assistant in the Teacher-Student Training Setup. | Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh |
| 2023 | ICASSP | Unfused: Unsupervised Finetuning Using Self Supervised Distillation. | Ashish Seth, Sreyan Ghosh, Srinivasan Umesh, Dinesh Manocha |
| 2023 | ICASSP | SLICER: Learning Universal Audio Representations Using Low-Resource Self-Supervised Pre-Training. | Ashish Seth, Sreyan Ghosh, Srinivasan Umesh, Dinesh Manocha |
| 2023 | ICASSP | Channel-Aware Pretraining Of Joint Encoder-Decoder Self-Supervised Model For Telephonic-Speech ASR. | Vrunda N. Sukhadia, Srinivasan Umesh |
| 2023 | Interspeech | Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages. | Anusha Prakash, Arun Kumar A, Ashish Seth, Bhagyashree Mukherjee, Ishika Gupta, Jom Kuriakose, Jordan Fernandes, K. V. Vikram, Mano Ranjith Kumar M., Metilda Sagaya Mary, Mohammad Wajahat, Mohana N, Mudit Batra, Navina K, Nihal John George, Nithya Ravi, Pruthwik Mishra, Sudhanshu Srivastava, Vasista Sai Lodagala, Vandan Mujadia, Kada Sai Venkata Vineeth, Vrunda N. Sukhadia, Dipti Misra Sharma, Hema A. Murthy, Pushpak Bhattacharyya, Srinivasan Umesh, Rajeev Sangal |
| 2023 | Interspeech | The Tag-Team Approach: Leveraging CLS and Language Tagging for Enhancing Multilingual ASR. | Kaousheik Jayakumar, Vrunda N. Sukhadia, Arun Kumar A, Srinivasan Umesh |
| 2023 | Interspeech | SALTTS: Leveraging Self-Supervised Speech Representations for improved Text-to-Speech Synthesis. | Ramanan Sivaguru, Vasista Sai Lodagala, Srinivasan Umesh |
| 2022 | ICASSP | Investigation of Robustness of Hubert Features from Different Layers to Domain, Accent and Language Variations. | Pratik Kumar, Vrunda N. Sukhadia, Srinivasan Umesh |
| 2022 | Interspeech | Investigation of Ensemble features of Self-Supervised Pretrained Models for Automatic Speech Recognition. | A. Arunkumar, Vrunda Nileshkumar Sukhadia, Srinivasan Umesh |
| 2022 | Interspeech | Joint Encoder-Decoder Self-Supervised Pre-training for ASR. | A. Arunkumar, Srinivasan Umesh |
| 2022 | Interspeech | Gram Vaani ASR Challenge on spontaneous telephone speech recordings in regional variations of Hindi. | Anish Bhanushali, Grant Bridgman, Deekshitha G, Prasanta Kumar Ghosh, Pratik Kumar, Saurabh Kumar, Adithya Raj Kolladath, Nithya Ravi, Aaditeshwar Seth, Ashish Seth, Abhayjeet Singh, Vrunda N. Sukhadia, Srinivasan Umesh, Sathvik Udupa, Lodagala V. S. V. Durga Prasad |
| 2022 | Interspeech | Span Classification with Structured Information for Disfluency Detection in Spoken Utterances. | Sreyan Ghosh, Sonal Kumar, Yaman Kumar, Rajiv Ratn Shah, Srinivasan Umesh |
| 2022 | Interspeech | DeToxy: A Large-Scale Multimodal Dataset for Toxicity Classification in Spoken Utterances. | Sreyan Ghosh, Samden Lepcha, Sakshi Singh, Rajiv Ratn Shah, Srinivasan Umesh |
| 2021 | ICASSP | Exploring the use of Common Label Set to Improve Speech Recognition of Low Resource Indian Languages. | Vishwas M. Shetty, Srinivasan Umesh |
| 2020 | ICASSP | Investigation of Methods to Improve the Recognition Performance of Tamil-English Code-Switched Data in Transformer Framework. | Metilda Sagaya Mary N. J, Vishwas M. Shetty, Srinivasan Umesh |
| 2020 | ICASSP | Improving the Performance of Transformer Based Low Resource Speech Recognition for Indian Languages. | Vishwas M. Shetty, Metilda Sagaya Mary N. J, Srinivasan Umesh |
| 2018 | Interspeech | Correlational Networks for Speaker Normalization in Automatic Speech Recognition. | Rini A. Sharon, Sandeep Reddy Kothinti, Srinivasan Umesh |
| 2018 | Interspeech | Articulatory and Stacked Bottleneck Features for Low Resource Speech Recognition. | Vishwas M. Shetty, Rini A. Sharon, Basil Abraham, Tejaswi Seeram, Anusha Prakash, Nithya Ravi, Srinivasan Umesh |
| 2018 | Interspeech | Investigating the Effect of Audio Duration on Dementia Detection Using Acoustic Features. | Jochen Weiner, Miguel Angrick, Srinivasan Umesh, Tanja Schultz |
| 2017 | Interspeech | Transfer Learning and Distillation Techniques to Improve the Acoustic Modeling of Low Resource Languages. | Basil Abraham, Tejaswi Seeram, Srinivasan Umesh |
| 2017 | Interspeech | Joint Estimation of Articulatory Features and Acoustic Models for Low-Resource Languages. | Basil Abraham, Srinivasan Umesh, Neethu Mariam Joy |
| 2017 | Interspeech | Generalized Distillation Framework for Speaker Normalization. | Neethu Mariam Joy, Sandeep Reddy Kothinti, Srinivasan Umesh, Basil Abraham |
| 2017 | Interspeech | On Improving Acoustic Models for TORGO Dysarthric Speech Database. | Neethu Mariam Joy, Srinivasan Umesh, Basil Abraham |
| 2016 | Interspeech | Articulatory Feature Extraction Using CTC to Build Articulatory Classifiers Without Forced Frame Alignments for Speech Recognition. | Basil Abraham, Srinivasan Umesh, Neethu Mariam Joy |
| 2016 | Interspeech | Overcoming Data Sparsity in Acoustic Modeling of Low-Resource Language by Borrowing Data and Model Parameters from High-Resource Languages. | Basil Abraham, Srinivasan Umesh, Neethu Mariam Joy |
| 2016 | Interspeech | DNNs for Unsupervised Extraction of Pseudo FMLLR Features Without Explicit Adaptation Data. | Neethu Mariam Joy, Murali Karthick Baskar, Srinivasan Umesh, Basil Abraham |
| 2015 | Interspeech | Speaker adaptation of convolutional neural network using speaker specific subspace vectors of SGMM. | Murali Karthick B, Prateek Kolhar, Srinivasan Umesh |
| 2013 | ASRU | Modified splice and its extension to non-stereo data for noise robust speech recognition. | D. S. Pavan Kumar, N. Vishnu Prasad, Vikas Joshi, Srinivasan Umesh |
| 2013 | ASRU | Acoustic modeling using transform-based phone-cluster adaptive training. | Vimal Manohar, Srinivas C. Bhargav, Srinivasan Umesh |
| 2013 | ASRU | Improved cepstral mean and variance normalization using Bayesian framework. | N. Vishnu Prasad, Srinivasan Umesh |
| 2013 | Interspeech | Modified cepstral mean normalization - transforming to utterance specific non-zero mean. | Vikas Joshi, N. Vishnu Prasad, Srinivasan Umesh |
| 2012 | ICASSP | Robust speech recognition through selection of speaker and environment transforms. | Raghavendra Bilgi, Vikas Joshi, Srinivasan Umesh, Luz Garca Martnez, M. Carmen Bentez Ortzar |
| 2012 | ICASSP | Noise and speaker compensation in the Log filter bank domain. | Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, Luz Garca Martnez, M. Carmen Bentez Ortzar |
| 2012 | ICASSP | Computationally efficient speaker identification using fast-MLLR based anchor modeling. | Achintya Kumar Sarkar, Srinivasan Umesh, Jean-Franois Bonastre |
| 2011 | ICASSP | Use of VTL-wise models in feature-mapping framework to achieve performance of multiple-background models in speaker verification. | Achintya Kumar Sarkar, Srinivasan Umesh |
| 2011 | Interspeech | Efficient Speaker and Noise Normalization for Robust Speech Recognition. | Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, M. Carmen Bentez, Luz Garca |
| 2011 | Interspeech | Sub-Band Level Histogram Equalization for Robust Speech Recognition. | Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, Luz Garca, M. Carmen Bentez |
| 2011 | Interspeech | Eigen-Voice Based Anchor Modeling System for Speaker Identification Using MLLR Super-Vector. | Achintya Kumar Sarkar, Srinivasan Umesh |
| 2010 | Interspeech | Fast computation of speaker characterization vector using MLLR and sufficient statistics in anchor model framework. | Achintya Kumar Sarkar, Srinivasan Umesh |
| 2009 | ICASSP | Improving the performance of VTLN under mismatched speaker conditions and making it approach that of matched speaker conditions. | D. Rama Sanand, Shakti Prasad Rath, Srinivasan Umesh |
| 2009 | Interspeech | Characterizing speaker variability using spectral envelopes of vowel sounds. | A. N. Harish, D. Rama Sanand, Srinivasan Umesh |
| 2009 | Interspeech | Acoustic class specific VTLN-warping using regression class trees. | Shakti Prasad Rath, Srinivasan Umesh |
| 2009 | Interspeech | Using VTLN matrices for rapid and computationally-efficient speaker adaptation with robustness to first-pass transcription errors. | Shakti Prasad Rath, Srinivasan Umesh, Achintya Kumar Sarkar |
| 2009 | Interspeech | A study on the influence of covariance adaptation on jacobian compensation in vocal tract length normalization. | D. Rama Sanand, Shakti Prasad Rath, Srinivasan Umesh |
| 2009 | Interspeech | Text-independent speaker identification using vocal tract length normalization for building universal background model. | Achintya Kumar Sarkar, Srinivasan Umesh, Shakti Prasad Rath |
| 2008 | Interspeech | A computationally efficient approach to warp factor estimation in VTLN using EM algorithm and sufficient statistics. | P. T. Akhil, Shakti Prasad Rath, Srinivasan Umesh, D. Rama Sanand |
| 2008 | Interspeech | Use of spectral centre of gravity for generating speaker invariant features for automatic speech recognition. | D. Rama Sanand, V. Balaji, Rani R. Sandhya, Srinivasan Umesh |
| 2008 | Interspeech | Study of jacobian compensation using linear transformation of conventional MFCC for VTLN. | D. Rama Sanand, Srinivasan Umesh |
| 2007 | IJCAI | Speaker-Invariant Features for Automatic Speech Recognition. | Srinivasan Umesh, D. Rama Sanand, G. Praveen |
| 2007 | Interspeech | Linear transformation approach to VTLN using dynamic frequency warping. | D. Rama Sanand, D. Dinesh Kumar, Srinivasan Umesh |
| 2006 | ICASSP | Study Of Non-Linear Frequency Warping Functions For Speaker Normalization. | S. V. Bharath Kumar, Srinivasan Umesh, Rohit Sinha |
| 2006 | ICASSP | Vtln Warping Factor Estimation Using Accumulation of Sufficient Statistics. | Jonas Lf, Hermann Ney, Srinivasan Umesh |
| 2005 | Interspeech | Implementing frequency-warping and VTLN through linear transformation of conventional MFCC. | Srinivasan Umesh, Andrs Zolnay, Hermann Ney |
| 2004 | ICASSP | Non-uniform speaker normalization using affine-transformation. | S. V. Bharath Kumar, Srinivasan Umesh, Rohit Sinha |
| 2004 | ICASSP | An investigation into front-end signal processing for speaker normalization. | Srinivasan Umesh, Rohit Sinha, S. V. Bharath Kumar |
| 2004 | Interspeech | Using VTLN for broadcast news transcription. | Do Yeong Kim, Srinivasan Umesh, Mark J. F. Gales, Thomas Hain, Philip C. Woodland |
| 2003 | ICASSP | A method for compensation of Jacobian in speaker normalization. | Rohit Sinha, Srinivasan Umesh |
| 2002 | ICASSP | Non-uniform scaling based speaker normalization. | Rohit Sinha, Srinivasan Umesh |
| 2002 | ICASSP | A simple approach to non-uniform vowel normalization. | Srinivasan Umesh, S. V. Bharath Kumar, M. K. Vinay, Rajesh Sharma, Rohit Sinha |
| 2000 | Interspeech | Exploiting frequency-scaling invariance properties of the scale transform for automatic speech recognition. | Srinivasan Umesh, Richard C. Rose, Sarangarajan Parthasarathy |
| 1999 | ICASSP | Fitting the Mel scale. | Srinivasan Umesh, Leon Cohen, Douglas J. Nelson |
| 1998 | ICASSP | Improved scale-cepstral analysis in speech. | Srinivasan Umesh, Leon Cohen, Douglas J. Nelson |
| 1997 | ICASSP | Frequency-warping and speaker-normalization. | Srinivasan Umesh, Leon Cohen, Douglas J. Nelson |
| 1996 | ICASSP | Computationally efficient estimation of sinusoidal frequency at low SNR. | Srinivasan Umesh, Douglas J. Nelson |
| 1996 | Interspeech | Frequency-warping in speech. | Srinivasan Umesh, Leon Cohen, Nenad Marinovic, Douglas J. Nelson |
| 1992 | ICASSP | Resolving the components of transient signals by a multistage procedure. | Srinivasan Umesh, Donald W. Tufts |