Kartik Audhkhasi
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
65
Venues
8
Active years
2007–2025
Best venue rank
A*
Where they publish
Papers
65 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | EMNLP | LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors. | Rao Ma, Tongzhou Chen, Kartik Audhkhasi, Bhuvana Ramabhadran |
| 2025 | ICASSP | Audio Diffusion with Large Language Models. | Yinghui Huang, Kyle Kastner, Kartik Audhkhasi, Bhuvana Ramabhadran, Andrew Rosenberg |
| 2025 | ICASSP | Weak-to-Strong Generalization in Speech Recognition. | Soheil Khorram, Qian Zhang, Rohit Prabhavalkar, Kartik Audhkhasi, Bhuvana Ramabhadran |
| 2025 | ICASSP | Identifying and Mitigating Mismatched Language Code in Multilingual ASR. | Jaeyoung Kim, Sepand Mavandadi, Kartik Audhkhasi, Shikhar Bharadwaj, Brian Farris, Tongzhou Chen, Bhuvana Ramabhadran, Sriram Ganapathy |
| 2024 | ICASSP | Task Vector Algebra for ASR Models. | Gowtham Ramesh, Kartik Audhkhasi, Bhuvana Ramabhadran |
| 2023 | ICASSP | Modular Conformer Training for Flexible End-to-End ASR. | Kartik Audhkhasi, Brian Farris, Bhuvana Ramabhadran, Pedro J. Moreno |
| 2023 | ICASSP | Large-Scale Language Model Rescoring on Long-Form Data. | Tongzhou Chen, Cyril Allauzen, Yinghui Huang, Daniel S. Park, David Rybach, W. Ronny Huang, Rodrigo Cabrera, Kartik Audhkhasi, Bhuvana Ramabhadran, Pedro J. Moreno, Michael Riley |
| 2023 | ICASSP | Robust Knowledge Distillation from RNN-T Models with Noisy Training Labels Using Full-Sum Loss. | Mohammad Zeineldeen, Kartik Audhkhasi, Murali Karthick Baskar, Bhuvana Ramabhadran |
| 2023 | Interspeech | O-1: Self-training with Oracle and 1-best Hypothesis. | Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran, Kartik Audhkhasi |
| 2022 | ACII | Federated Learning for Affective Computing Tasks. | Krishna Somandepalli, Hang Qi, Brian Eoff, Alan Cowen, Kartik Audhkhasi, Josh Belanich, Brendan Jou |
| 2022 | Interspeech | Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition. | Kartik Audhkhasi, Yinghui Huang, Bhuvana Ramabhadran, Pedro J. Moreno |
| 2021 | ICASSP | Convolutional Dropout and Wordpiece Augmentation for End-to-End Speech Recognition. | Hainan Xu, Yinghui Huang, Yun Zhu, Kartik Audhkhasi, Bhuvana Ramabhadran |
| 2021 | Interspeech | Mixture Model Attention: Flexible Streaming and Non-Streaming Automatic Speech Recognition. | Kartik Audhkhasi, Tongzhou Chen, Bhuvana Ramabhadran, Pedro J. Moreno |
| 2021 | Interspeech | AVLnet: Learning Audio-Visual Language Representations from Instructional Videos. | Andrew Rouditchenko, Angie W. Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogrio Schmidt Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, James R. Glass |
| 2021 | Interspeech | Regularizing Word Segmentation by Creating Misspellings. | Hainan Xu, Kartik Audhkhasi, Yinghui Huang, Jesse Emond, Bhuvana Ramabhadran |
| 2020 | ICASSP | Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent Systems. | Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas, Zvi Kons, Kartik Audhkhasi, Brian Kingsbury, Ron Hoory, Michael Picheny |
| 2020 | ICASSP | Alignment-Length Synchronous Decoding for RNN Transducer. | George Saon, Zoltn Tske, Kartik Audhkhasi |
| 2020 | Interspeech | Transliteration Based Data Augmentation for Training Multilingual ASR Acoustic Models in Low Resource Settings. | Samuel Thomas, Kartik Audhkhasi, Brian Kingsbury |
| 2020 | Interspeech | End-to-End Spoken Language Understanding Without Full Transcripts. | Hong-Kwang Jeff Kuo, Zoltn Tske, Samuel Thomas, Yinghui Huang, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, Luis A. Lastras |
| 2020 | Interspeech | Single Headed Attention Based Sequence-to-Sequence Model for State-of-the-Art Results on Switchboard. | Zoltn Tske, George Saon, Kartik Audhkhasi, Brian Kingsbury |
| 2019 | ASRU | Simplified LSTMS for Speech Recognition. | George Saon, Zoltn Tske, Kartik Audhkhasi, Brian Kingsbury, Michael Picheny, Samuel Thomas |
| 2019 | CVPR | Grounding Spoken Words in Unlabeled Video. | Angie W. Boggust, Kartik Audhkhasi, Dhiraj Joshi, David Harwath, Samuel Thomas, Rogrio Schmidt Feris, Danny Gutfreund, Yang Zhang, Antonio Torralba, Michael Picheny, James R. Glass |
| 2019 | ICASSP | Sequence Noise Injected Training for End-to-end Speech Recognition. | George Saon, Zoltn Tske, Kartik Audhkhasi, Brian Kingsbury |
| 2019 | ICASSP | Acoustically Grounded Word Embeddings for Improved Acoustics-to-word Speech Recognition. | Shane Settle, Kartik Audhkhasi, Karen Livescu, Michael Picheny |
| 2019 | Interspeech | Forget a Bit to Learn Better: Soft Forgetting for CTC-Based Automatic Speech Recognition. | Kartik Audhkhasi, George Saon, Zoltn Tske, Brian Kingsbury, Michael Picheny |
| 2019 | Interspeech | Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge Distillation. | Gakuto Kurata, Kartik Audhkhasi |
| 2019 | Interspeech | Multi-Task CTC Training with Auxiliary Feature Reconstruction for End-to-End Speech Recognition. | Gakuto Kurata, Kartik Audhkhasi |
| 2019 | Interspeech | Challenging the Boundaries of Speech Recognition: The MALACH Corpus. | Michael Picheny, Zoltn Tske, Brian Kingsbury, Kartik Audhkhasi, Xiaodong Cui, George Saon |
| 2019 | Interspeech | Detection and Recovery of OOVs for Improved English Broadcast News Captioning. | Samuel Thomas, Kartik Audhkhasi, Zoltn Tske, Yinghui Huang, Michael Picheny |
| 2019 | Interspeech | Advancing Sequence-to-Sequence Based Speech Recognition. | Zoltn Tske, Kartik Audhkhasi, George Saon |
| 2018 | ICASSP | Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech Recognition. | Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Michael Picheny |
| 2018 | ICASSP | Whole Sentence Neural Language Models. | Yinghui Huang, Abhinav Sethy, Kartik Audhkhasi, Bhuvana Ramabhadran |
| 2018 | ICASSP | Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition. | Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg, Samuel Thomas, Bhuvana Ramabhadran, Mark Hasegawa-Johnson |
| 2017 | ICASSP | End-to-end ASR-free keyword search from speech. | Kartik Audhkhasi, Andrew Rosenberg, Abhinav Sethy, Bhuvana Ramabhadran, Brian Kingsbury |
| 2017 | ICASSP | Knowledge distillation across ensembles of multilingual models for low-resource languages. | Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Tom Sercu, Kartik Audhkhasi, Abhinav Sethy, Markus Nubaum-Thom, Andrew Rosenberg |
| 2017 | ICASSP | End-to-end speech recognition and keyword search on low-resource languages. | Andrew Rosenberg, Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Michael Picheny |
| 2017 | Interspeech | Direct Acoustics-to-Word Models for English Conversational Speech Recognition. | Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon, Michael Picheny, David Nahamoo |
| 2017 | Interspeech | English Conversational Telephone Speech Recognition by Humans and Machines. | George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas, Dimitrios Dimitriadis, Xiaodong Cui, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
| 2016 | ICASSP | Semantic word embedding neural network language models for automatic speech recognition. | Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran |
| 2016 | ICASSP | Efficient one-vs-one kernel ridge regression for speech recognition. | Jie Chen, Lingfei Wu, Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran |
| 2016 | Interspeech | Multilingual Data Selection for Low Resource Speech Recognition. | Samuel Thomas, Kartik Audhkhasi, Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran |
| 2015 | ASRU | Multilingual representations for low resource speech recognition and keyword search. | Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Xiaodong Cui, Ellen Kislal, Lidia Mangu, Markus Nubaum-Thom, Michael Picheny, Zoltn Tske, Pavel Golik, Ralf Schlter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Haipeng Wang, Philip C. Woodland |
| 2015 | ICASSP | A mixture of experts approach towards intelligibility classification of pathological speech. | Rahul Gupta, Kartik Audhkhasi, Shrikanth S. Narayanan |
| 2014 | ICASSP | Fusion of diverse denoising systems for robust automatic speech recognition. | Naveen Kumar, Maarten Van Segbroeck, Kartik Audhkhasi, Peter Drotr, Shrikanth S. Narayanan |
| 2014 | ICASSP | Semi-supervised term-weighted value rescoring for keyword search. | Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Shrikanth S. Narayanan |
| 2014 | ICASSP | Training ensemble of diverse classifiers on feature subsets. | Rahul Gupta, Kartik Audhkhasi, Shrikanth S. Narayanan |
| 2013 | ASRU | Joint training of interpolated exponential n-gram models. | Abhinav Sethy, Stanley F. Chen, Ebru Arisoy, Bhuvana Ramabhadran, Kartik Audhkhasi, Shrikanth S. Narayanan, Paul Vozila |
| 2013 | IJCNN | Noise benefits in backpropagation and deep bidirectional pre-training. | Kartik Audhkhasi, Osonde Osoba, Bart Kosko |
| 2013 | IJCNN | Noisy hidden Markov models for speech recognition. | Kartik Audhkhasi, Osonde Osoba, Bart Kosko |
| 2013 | Interspeech | Empirical link between hypothesis diversity and fusion performance in an ensemble of automatic speech recognition systems. | Kartik Audhkhasi, Andreas M. Zavou, Panayiotis G. Georgiou, Shrikanth S. Narayanan |
| 2013 | Interspeech | Classifying language-related developmental disorders from speech cues: the promise and the potential confounds. | Daniel Bone, Theodora Chaspari, Kartik Audhkhasi, James Gibson, Andreas Tsiartas, Maarten Van Segbroeck, Ming Li, Sungbok Lee, Shrikanth S. Narayanan |
| 2013 | Interspeech | Paralinguistic event detection from speech using probabilistic time-series smoothing and masking. | Rahul Gupta, Kartik Audhkhasi, Sungbok Lee, Shrikanth S. Narayanan |
| 2013 | SIGdial | Which ASR should I choose for my dialogue system? | Fabrizio Morbini, Kartik Audhkhasi, Kenji Sagae, Ron Artstein, Dogan Can, Panayiotis G. Georgiou, Shrikanth S. Narayanan, Anton Leuski, David R. Traum |
| 2012 | ICASSP | Analyzing quality of crowd-sourced speech transcriptions of noisy audio for acoustic model adaptation. | Kartik Audhkhasi, Panayiotis G. Georgiou, Shrikanth S. Narayanan |
| 2012 | ICASSP | Creating ensemble of diverse maximum entropy models. | Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Shrikanth S. Narayanan |
| 2012 | Interspeech | Speaker Personality Classification Using Systems Based on Acoustic-Lexical Cues and an Optimal Tree-Structured Bayesian Network. | Kartik Audhkhasi, Angeliki Metallinou, Ming Li, Shrikanth S. Narayanan |
| 2011 | ICASSP | Accurate transcription of broadcast news speech using multiple noisy transcribers and unsupervised reliability metrics. | Kartik Audhkhasi, Panayiotis G. Georgiou, Shrikanth S. Narayanan |
| 2011 | ICASSP | Emotion classification from speech using evaluator reliability-weighted combination of ranked lists. | Kartik Audhkhasi, Shrikanth S. Narayanan |
| 2011 | Interspeech | Reliability-Weighted Acoustic Model Adaptation Using Crowd Sourced Transcriptions. | Kartik Audhkhasi, Panayiotis G. Georgiou, Shrikanth S. Narayanan |
| 2010 | Interspeech | Data-dependent evaluator modeling and its application to emotional valence classification from speech. | Kartik Audhkhasi, Shrikanth S. Narayanan |
| 2010 | Interspeech | Automatic speech recognition system channel modeling. | Qun Feng Tan, Kartik Audhkhasi, Panayiotis G. Georgiou, Emil Ettelaie, Shrikanth S. Narayanan |
| 2009 | ASRU | Lattice-based lexical cues for word fragment detection in conversational speech. | Kartik Audhkhasi, Panayiotis G. Georgiou, Shrikanth S. Narayanan |
| 2009 | ICASSP | Formant-based technique for automatic filled-pause detection in spontaneous spoken english. | Kartik Audhkhasi, Kundan Kandhway, Om Deshmukh, Ashish Verma |
| 2009 | ICASSP | Automatic evaluation of spoken english fluency. | Om Deshmukh, Kundan Kandhway, Ashish Verma, Kartik Audhkhasi |
| 2007 | ICASSP | Keyword Search using Modified Minimum Edit Distance Measure. | Kartik Audhkhasi, Ashish Verma |