| 2025 | AAAI | Audio Entailment: Assessing Deductive Reasoning for Audio Understanding. | Soham Deshmukh, Shuo Han, Hazim T. Bukhari, Benjamin Elizalde, Hannes Gamper, Rita Singh, Bhiksha Raj |
| 2025 | ACL | Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models. | Hanin Atwany, Abdul Waheed, Rita Singh, Monojit Choudhury, Bhiksha Raj |
| 2025 | ACL | On the Robust Approximation of ASR Metrics. | Abdul Waheed, Hanin Atwany, Rita Singh, Bhiksha Raj |
| 2025 | ASRU | CoLMbo: Speaker Language Model for Descriptive Profiling. | Massa Baali, Shuo Han, Syed Abdul Hannan, Purusottam Samal, Karanveer Singh, Soham Deshmukh, Rita Singh, Bhiksha Raj |
| 2025 | CIKM | PlaceSim: An LLM-based Interactive Platform for Human Behavior Simulation in Physical Facilities. | Suhyeon Lee, Youngjun Yu, Donghyuk Shin, Rita Singh |
| 2025 | EMNLP | SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions. | Massa Baali, Sarthak Bisht, Francisco Teixeira, Kateryna Shapovalenko, Rita Singh, Bhiksha Raj |
| 2025 | EMNLP | CAARMA: Class Augmentation with Adversarial Mixup Regularization. | Massa Baali, Xiang Li, Hao Chen, Syed Abdul Hannan, Rita Singh, Bhiksha Raj |
| 2025 | EMNLP | PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs. | Sana Kang, Myeongseok Gwon, Su Young Kwon, Jaewook Lee, Andrew Lan, Bhiksha Raj, Rita Singh |
| 2025 | ICASSP | Tessellated Linear Model for Age Prediction from Voice. | Dareen Alharthi, Mahsa Zamani, Bhiksha Raj, Rita Singh |
| 2025 | ICCV | Does Prior Data Matter? Exploring Joint Training in the Context of Few-Shot Class-Incremental Learning. | Shiwon Kim, Dongjun Hwang, Sungwon Woo, Rita Singh |
| 2025 | ICLR | ADIFF: Explaining audio difference using natural language. | Soham Deshmukh, Shuo Han, Rita Singh, Bhiksha Raj |
| 2024 | CVPR | QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition. | Xiang Li, Jinglu Wang, Xiaohao Xu, Xiulian Peng, Rita Singh, Yan Lu, Bhiksha Raj |
| 2024 | ECCV | R | Xiang Li, Kai Qiu, Jinglu Wang, Xiaohao Xu, Rita Singh, Kashu Yamazaki, Hao Chen, Xiaonan Huang, Bhiksha Raj |
| 2024 | ICASSP | Importance of Negative Sampling in Weak Label Learning. | Ankit Shah, Fuyu Tang, Zelin Ye, Rita Singh, Bhiksha Raj |
| 2024 | ICASSP | Training Audio Captioning Models without Audio. | Soham Deshmukh, Benjamin Elizalde, Dimitra Emmanouilidou, Bhiksha Raj, Rita Singh, Huaming Wang |
| 2024 | ICASSP | Prompting Audios Using Acoustic Properties for Emotion Representation. | Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh, Huaming Wang, Bhiksha Raj, Rita Singh |
| 2024 | ICASSP | Vocal Fold Dynamics for Automatic Detection of Amyotrophic Lateral Sclerosis from Voice. | Jiayi Zhang, Rita Singh |
| 2024 | ICML | A General Framework for Learning from Weak Supervision. | Hao Chen, Jindong Wang, Lei Feng, Xiang Li, Yidong Wang, Xing Xie, Masashi Sugiyama, Rita Singh, Bhiksha Raj |
| 2024 | ICML | Completing Visual Objects via Bridging Generation and Segmentation. | Xiang Li, Yinpeng Chen, Chung-Ching Lin, Hao Chen, Kai Hu, Rita Singh, Bhiksha Raj, Lijuan Wang, Zicheng Liu |
| 2024 | Interspeech | SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios. | Hazim T. Bukhari, Soham Deshmukh, Hira Dhamyal, Bhiksha Raj, Rita Singh |
| 2024 | Interspeech | PAM: Prompting Audio-Language Models for Audio Quality Assessment. | Soham Deshmukh, Dareen Alharthi, Benjamin Elizalde, Hannes Gamper, Mahmoud Al Ismail, Rita Singh, Bhiksha Raj, Huaming Wang |
| 2024 | Interspeech | Domain Adaptation for Contrastive Audio-Language Models. | Soham Deshmukh, Rita Singh, Bhiksha Raj |
| 2024 | NAACL | R-BASS : Relevance-aided Block-wise Adaptation for Speech Summarization. | Roshan Sharma, Ruchira Sharma, Hira Dhamyal, Rita Singh, Bhiksha Raj |
| 2023 | ASRU | Espnet-Summ: Introducing a Novel Large Dataset, Toolkit, and a Cross-Corpora Evaluation of Speech Summarization Systems. | Roshan S. Sharma, William Chen, Takatomo Kano, Ruchira Sharma, Siddhant Arora, Shinji Watanabe, Atsunori Ogawa, Marc Delcroix, Rita Singh, Bhiksha Raj |
| 2023 | EMNLP | Token Prediction as Implicit Classification to Identify LLM-Generated Text. | Yutian Chen, Hao Kang, Vivian Zhai, Liangze Li, Rita Singh, Bhiksha Raj |
| 2023 | EMNLP | Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and Text. | Xiang Li, Jinglu Wang, Xiaohao Xu, Muqiao Yang, Fan Yang, Yizhou Zhao, Rita Singh, Bhiksha Raj |
| 2023 | ICCV | Pairwise Similarity Learning is SimPLE. | Yandong Wen, Weiyang Liu, Yao Feng, Bhiksha Raj, Rita Singh, Adrian Weller, Michael J. Black, Bernhard Schlkopf |
| 2023 | Interspeech | BASS: Block-wise Adaptation for Speech Summarization. | Roshan Sharma, Siddhant Arora, Kenneth Zheng, Shinji Watanabe, Rita Singh, Bhiksha Raj |
| 2023 | Interspeech | The Hidden Dance of Phonemes and Visage: Unveiling the Enigmatic Link between Phonemes and Facial Features. | Liao Qu, Xianwei Zou, Xiang Li, Yandong Wen, Rita Singh, Bhiksha Raj |
| 2022 | ICLR | SphereFace2: Binary Classification is All You Need for Deep Face Recognition. | Yandong Wen, Weiyang Liu, Adrian Weller, Bhiksha Raj, Rita Singh |
| 2022 | Interspeech | Positional Encoding for Capturing Modality Specific Cadence for Emotion Detection. | Hira Dhamyal, Bhiksha Raj, Rita Singh |
| 2021 | ICASSP | Interpreting Glottal Flow Dynamics for Detecting Covid-19 From Voice. | Soham Deshmukh, Mahmoud Al Ismail, Rita Singh |
| 2021 | ICASSP | Detection of Covid-19 Through the Analysis of Vocal Fold Oscillations. | Mahmoud Al Ismail, Soham Deshmukh, Rita Singh |
| 2021 | ICCV | Self-Supervised 3D Face Reconstruction via Conditional Estimation. | Yandong Wen, Weiyang Liu, Bhiksha Raj, Rita Singh |
| 2021 | Interspeech | Improving Weakly Supervised Sound Event Detection with Self-Supervised Auxiliary Tasks. | Soham Deshmukh, Bhiksha Raj, Rita Singh |
| 2021 | Interspeech | Generalized Spoofing Detection Inspired from Audio Generation Artifacts. | Yang Gao, Tyler Vuong, Mahsa Elyasi, Gaurav Bharaj, Rita Singh |
| 2021 | Interspeech | Masked Proxy Loss for Text-Independent Speaker Verification. | Jiachen Lian, Aiswarya Vinod Kumar, Hira Dhamyal, Bhiksha Raj, Rita Singh |
| 2020 | ICASSP | Speech-Based Parameter Estimation of an Asymmetric Vocal Fold Oscillation Model and its Application in Discriminating Vocal Fold Pathologies. | Wenbo Zhao, Rita Singh |
| 2020 | ICPR | Hierarchical Routing Mixture of Experts. | Wenbo Zhao, Yang Gao, Shahan Ali Memon, Bhiksha Raj, Rita Singh |
| 2020 | Interspeech | The Phonetic Bases of Vocal Expressed Emotion: Natural versus Acted. | Hira Dhamyal, Shahan Ali Memon, Bhiksha Raj, Rita Singh |
| 2020 | Interspeech | Hide and Speak: Towards Deep Neural Networks for Speech Steganography. | Felix Kreuk, Yossi Adi, Bhiksha Raj, Rita Singh, Joseph Keshet |
| 2020 | ISVC | Controlled AutoEncoders to Generate Faces from Voices. | Hao Liang, Lulan Yu, Guikang Xu, Bhiksha Raj, Rita Singh |
| 2019 | ASRU | Optimizing Neural Network Embeddings Using a Pair-Wise Loss for Text-Independent Speaker Verification. | Hira Dhamyal, Tianyan Zhou, Bhiksha Raj, Rita Singh |
| 2019 | ICASSP | Human Behaviour Recognition Using Wifi Channel State Information. | Daanish Ali Khan, Saquib Razak, Bhiksha Raj, Rita Singh |
| 2019 | ICLR | Disjoint Mapping Network for Cross-modal Matching of Voices and Faces. | Yandong Wen, Mahmoud Al Ismail, Weiyang Liu, Bhiksha Raj, Rita Singh |
| 2019 | IJCNN | Neural Regression Trees. | Shahan Ali Memon, Wenbo Zhao, Bhiksha Raj, Rita Singh |
| 2018 | ICASSP | Voice Impersonation Using Generative Adversarial Networks. | Yang Gao, Rita Singh, Bhiksha Raj |
| 2018 | ICASSP | A Corrective Learning Approach for Text-Independent Speaker Verification. | Yandong Wen, Tianyan Zhou, Rita Singh, Bhiksha Raj |
| 2017 | ICASSP | Supervised monaural source separation based on autoencoders. | Keiichi Osako, Yuki Mitsufuji, Rita Singh, Bhiksha Raj |
| 2016 | ICASSP | The relationship of voice onset time and Voice Offset Time to physical age. | Rita Singh, Joseph Keshet, Deniz Genaga, Bhiksha Raj |
| 2016 | Interspeech | Estimation of Children's Physical Characteristics from Their Voices. | Jill Fain Lehman, Rita Singh |
| 2015 | ICASSP | Free energy for speech recognition. | Rita Singh, Ken'ichi Kumatani |
| 2015 | Interspeech | Keyword spotting in multi-player voice driven games for children. | Sundar Harshavardhan, Jill Fain Lehman, Rita Singh |
| 2014 | IC2E | Audio Classification with Thermodynamic Criteria. | Rita Singh |
| 2013 | ICASSP | Joint constrained maximum likelihood regression for overlapping speech recognition. | Ken'ichi Kumatani, Rita Singh, Friedrich Faubel, John W. McDonough, Youssef Oualil |
| 2013 | Interspeech | Discriminatively trained dependency language modeling for conversational speech recognition. | Benjamin Lambert, Bhiksha Raj, Rita Singh |
| 2012 | ICASSP | Spectrographic seam patterns for discriminative word spotting. | Shubhranshu Barnwal, Kamal Sahni, Rita Singh, Bhiksha Raj |
| 2012 | ICASSP | Audio event detection from acoustic unit occurrence patterns. | Anurag Kumar, Pranay Dighe, Rita Singh, Sourish Chaudhuri, Bhiksha Raj |
| 2012 | ICASSP | Compensating for denoising artifacts. | Rita Singh |
| 2012 | Interspeech | Exploiting Temporal Sequence Structure for Semantic Analysis of Multimedia. | Sourish Chaudhuri, Rita Singh, Bhiksha Raj |
| 2012 | Interspeech | Plagiarism Detection in Polyphonic Music using Monaural Signal Separation. | Soham De, Indradyumna Roy, Tarunima Prabhakar, Kriti Suneja, Sourish Chaudhuri, Rita Singh, Bhiksha Raj |
| 2012 | Interspeech | Microphone Array Post-filter based on Spatially-Correlated Noise Measurements for Distant Speech Recognition. | Ken'ichi Kumatani, Bhiksha Raj, Rita Singh, John W. McDonough |
| 2012 | Interspeech | Language identification using spectro-temporal patch features. | Kamal Sahni, Pranay Dighe, Rita Singh, Bhiksha Raj |
| 2012 | Interspeech | A signal-separation-based array postfilter for distant speech recognition. | Rita Singh, Ken'ichi Kumatani, John W. McDonough, Chen Liu |
| 2011 | ICASSP | An iterative least-squares technique for dereverberation. | Kshitiz Kumar, Bhiksha Raj, Rita Singh, Richard M. Stern |
| 2011 | ICASSP | Gammatone sub-band magnitude-domain dereverberation for ASR. | Kshitiz Kumar, Rita Singh, Bhiksha Raj, Richard M. Stern |
| 2011 | ICASSP | A paired test for recognizer selection with untranscribed data. | Bhiksha Raj, Rita Singh, James Baker |
| 2011 | Interspeech | Phoneme-Dependent NMF for Speech Enhancement in Monaural Mixtures. | Bhiksha Raj, Rita Singh, Tuomas Virtanen |
| 2010 | ICASSP | Latent-variable decomposition based dereverberation of monaural and multi-channel signals. | Rita Singh, Bhiksha Raj, Paris Smaragdis |
| 2010 | Interspeech | Creating a linguistic plausibility dataset with non-expert annotators. | Benjamin Lambert, Rita Singh, Bhiksha Raj |
| 2010 | Interspeech | Non-negative matrix factorization based compensation of music for automatic speech recognition. | Bhiksha Raj, Tuomas Virtanen, Sourish Chaudhuri, Rita Singh |
| 2010 | Interspeech | The use of sense in unsupervised training of acoustic models for ASR systems. | Rita Singh, Benjamin Lambert, Bhiksha Raj |
| 2009 | ICASSP | A joint decoding algorithm for multiple-example-based addition of words to a pronunciation lexicon. | Dhananjay Bansal, Nishanth Ulhas Nair, Rita Singh, Bhiksha Raj |
| 2007 | ICASSP | Bandwidth Expansionwith a plya URN Model. | Bhiksha Raj, Rita Singh, Madhusudana V. S. Shashanka, Paris Smaragdis |
| 2007 | Interspeech | Probabilistic deduction of symbol mappings for extension of lexicons. | Rita Singh, Evandro B. Gouva, Bhiksha Raj |
| 2005 | Interspeech | Recognizing speech from simultaneous speakers. | Bhiksha Raj, Rita Singh, Paris Smaragdis |
| 2004 | ICASSP | On tracking noise with linear dynamical system models. | Bhiksha Raj, Rita Singh, Richard M. Stern |
| 2004 | Interspeech | Maximum - likelihod adaptation of semi-continuous HMMs by latent variable decomposition of state distributions. | Antoine Raux, Rita Singh |
| 2003 | ICASSP | Tracking noise via dynamical systems with a continuum of states. | Rita Singh, Bhiksha Raj |
| 2003 | Interspeech | Design of the CMU sphinx-4 decoder. | Paul Lamere, Philip Kwok, William Walker, Evandro B. Gouva, Rita Singh, Bhiksha Raj, Peter Wolf |
| 2003 | Interspeech | Classification with free energy at raised temperatures. | Rita Singh, Manfred K. Warmuth, Bhiksha Raj, Paul Lamere |
| 2002 | Interspeech | Rapid development of speech-to-speech translation systems. | Alan W. Black, Ralf D. Brown, Robert E. Frederking, Kevin A. Lenzo, John Moody, Alexander I. Rudnicky, Rita Singh, Eric Steinbrecher |
| 2002 | Interspeech | Combining search spaces of heterogeneous recognizers for improved speech recogniton. | Xiang Li, Rita Singh, Richard M. Stern |
| 2001 | ICASSP | Tandem acoustic modeling in large-vocabulary recognition. | Daniel P. W. Ellis, Rita Singh, Sunil Sivadas |
| 2001 | ICASSP | Speech in Noisy Environments: robust automatic segmentation, feature extraction, and hypothesis combination. | Rita Singh, Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
| 2000 | ICASSP | Automatic generation of phone sets and lexical transcriptions. | Rita Singh, Bhiksha Raj, Richard M. Stern |
| 2000 | Interspeech | Phone transition acoustic modeling: application to speaker independent and spontaneous speech systems. | Jon P. Nedel, Rita Singh, Richard M. Stern |
| 2000 | Interspeech | Automatic subword unit refinement for spontaneous speech recognition via phone splitting. | Jon P. Nedel, Rita Singh, Richard M. Stern |
| 2000 | Interspeech | Task and domain specific modelling in the Carnegie Mellon communicator system. | Alexander I. Rudnicky, Christina L. Bennett, Alan W. Black, Ananlada Chotimongkol, Kevin A. Lenzo, Alice Oh, Rita Singh |
| 2000 | Interspeech | Structured redefinition of sound units by merging and splitting for improved speech recognition. | Rita Singh, Bhiksha Raj, Richard M. Stern |
| 1999 | ICASSP | Automatic clustering and generation of contextual questions for tied states in hidden Markov models. | Rita Singh, Bhiksha Raj, Richard M. Stern |
| 1999 | Interspeech | Domain adduced state tying for cross-domain acoustic modelling. | Rita Singh, Bhiksha Raj, Richard M. Stern |
| 1998 | Interspeech | Inference of missing spectrographic features for robust speech recognition. | Bhiksha Raj, Rita Singh, Richard M. Stern |