| 2025 | AAAI | Audio Entailment: Assessing Deductive Reasoning for Audio Understanding. | Soham Deshmukh, Shuo Han, Hazim T. Bukhari, Benjamin Elizalde, Hannes Gamper, Rita Singh, Bhiksha Raj |
| 2025 | ACL | Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models. | Hanin Atwany, Abdul Waheed, Rita Singh, Monojit Choudhury, Bhiksha Raj |
| 2025 | ACL | On the Robust Approximation of ASR Metrics. | Abdul Waheed, Hanin Atwany, Rita Singh, Bhiksha Raj |
| 2025 | ASRU | CoLMbo: Speaker Language Model for Descriptive Profiling. | Massa Baali, Shuo Han, Syed Abdul Hannan, Purusottam Samal, Karanveer Singh, Soham Deshmukh, Rita Singh, Bhiksha Raj |
| 2025 | CVPR | SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer. | Hao Chen, Ze Wang, Xiang Li, Ximeng Sun, Fangyi Chen, Jiang Liu, Jindong Wang, Bhiksha Raj, Zicheng Liu, Emad Barsoum |
| 2025 | CVPR | FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene Understanding. | Thanh-Dat Truong, Utsav Prabhu, Bhiksha Raj, Jackson David Cothren, Khoa Luu |
| 2025 | EMNLP | SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions. | Massa Baali, Sarthak Bisht, Francisco Teixeira, Kateryna Shapovalenko, Rita Singh, Bhiksha Raj |
| 2025 | EMNLP | CAARMA: Class Augmentation with Adversarial Mixup Regularization. | Massa Baali, Xiang Li, Hao Chen, Syed Abdul Hannan, Rita Singh, Bhiksha Raj |
| 2025 | EMNLP | PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs. | Sana Kang, Myeongseok Gwon, Su Young Kwon, Jaewook Lee, Andrew Lan, Bhiksha Raj, Rita Singh |
| 2025 | ICASSP | Tessellated Linear Model for Age Prediction from Voice. | Dareen Alharthi, Mahsa Zamani, Bhiksha Raj, Rita Singh |
| 2025 | ICASSP | MACE: Leveraging Audio for Evaluating Audio Captioning Systems. | Satvik Dixit, Soham Deshmukh, Bhiksha Raj |
| 2025 | ICASSP | Revisiting Acoustic Features for Robust ASR. | Muhammad A. Shah, Bhiksha Raj |
| 2025 | ICCV | Toward Material-Agnostic System Identification From Videos. | Yizhou Zhao, Haoyu Chen, Chunjiang Liu, Zhenyang Li, Charles Herrmann, Junhwa Hur, Yinxiao Li, Ming-Hsuan Yang, Bhiksha Raj, Min Xu |
| 2025 | ICCV | Total-Editing: Head Avatar with Editable Appearance, Motion, and Lighting. | Yizhou Zhao, Chunjiang Liu, Haoyu Chen, Bhiksha Raj, Min Xu, Tadas Baltrusaitis, Mitch Rundle, HsiangTao Wu, Kamran Ghasedi |
| 2025 | ICLR | ImageFolder: Autoregressive Image Generation with Folded Tokens. | Xiang Li, Kai Qiu, Hao Chen, Jason Kuen, Jiuxiang Gu, Bhiksha Raj, Zhe Lin |
| 2025 | ICLR | ADIFF: Explaining audio difference using natural language. | Soham Deshmukh, Shuo Han, Rita Singh, Bhiksha Raj |
| 2025 | ICLR | Speech Robust Bench: A Robustness Benchmark For Speech Recognition. | Muhammad A. Shah, David Solans Noguero, Mikko A. Heikkil, Bhiksha Raj, Nicolas Kourtellis |
| 2025 | ICLR | Unsupervised Disentanglement of Content and Style via Variance-Invariance Constraints. | Yuxuan Wu, Ziyu Wang, Bhiksha Raj, Gus Xia |
| 2025 | ICLR | Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video. | Xiaohao Xu, Tianyi Zhang, Shibo Zhao, Xiang Li, Sibo Wang, Yongqi Chen, Ye Li, Bhiksha Raj, Matthew Johnson-Roberson, Sebastian A. Scherer, Xiaonan Huang |
| 2025 | ICML | Masked Autoencoders Are Effective Tokenizers for Diffusion Models. | Hao Chen, Yujin Han, Fangyi Chen, Xiang Li, Yidong Wang, Jindong Wang, Ze Wang, Zicheng Liu, Difan Zou, Bhiksha Raj |
| 2025 | NAACL | uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes. | Abdul Waheed, Karima Kadaoui, Bhiksha Raj, Muhammad Abdul-Mageed |
| 2024 | ACL | Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization? | Roshan Sharma, Suwon Shon, Mark Lindsey, Hira Dhamyal, Bhiksha Raj |
| 2024 | ACL | Continual Contrastive Spoken Language Understanding. | Umberto Cappellazzo, Enrico Fini, Muqiao Yang, Daniele Falavigna, Alessio Brutti, Bhiksha Raj |
| 2024 | CVPR | QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition. | Xiang Li, Jinglu Wang, Xiaohao Xu, Xiulian Peng, Rita Singh, Yan Lu, Bhiksha Raj |
| 2024 | CVPR | Synergistic Global-Space Camera and Human Reconstruction from Videos. | Yizhou Zhao, Tuanfeng Yang Wang, Bhiksha Raj, Min Xu, Jimei Yang, Chun-Hao Paul Huang |
| 2024 | ECCV | R | Xiang Li, Kai Qiu, Jinglu Wang, Xiaohao Xu, Rita Singh, Kashu Yamazaki, Hao Chen, Xiaonan Huang, Bhiksha Raj |
| 2024 | ICASSP | Importance of Negative Sampling in Weak Label Learning. | Ankit Shah, Fuyu Tang, Zelin Ye, Rita Singh, Bhiksha Raj |
| 2024 | ICASSP | Training Audio Captioning Models without Audio. | Soham Deshmukh, Benjamin Elizalde, Dimitra Emmanouilidou, Bhiksha Raj, Rita Singh, Huaming Wang |
| 2024 | ICASSP | Prompting Audios Using Acoustic Properties for Emotion Representation. | Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh, Huaming Wang, Bhiksha Raj, Rita Singh |
| 2024 | ICASSP | AugSumm: Towards Generalizable Speech Summarization Using Synthetic Labels from Large Language Models. | Jee-Weon Jung, Roshan S. Sharma, William Chen, Bhiksha Raj, Shinji Watanabe |
| 2024 | ICASSP | Fixed Inter-Neuron Covariability Induces Adversarial Robustness. | Muhammad A. Shah, Bhiksha Raj |
| 2024 | ICASSP | Improving Continual Learning of Acoustic Scene Classification via Mutual Information Optimization. | Muqiao Yang, Umberto Cappellazzo, Xiang Li, Bhiksha Raj |
| 2024 | ICASSP | uSee: Unified Speech Enhancement And Editing with Conditional Diffusion Models. | Muqiao Yang, Chunlei Zhang, Yong Xu, Zhongweiyang Xu, Heming Wang, Bhiksha Raj, Dong Yu |
| 2024 | ICLR | Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks. | Hao Chen, Jindong Wang, Ankit Shah, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, Bhiksha Raj |
| 2024 | ICML | A General Framework for Learning from Weak Supervision. | Hao Chen, Jindong Wang, Lei Feng, Xiang Li, Yidong Wang, Xing Xie, Masashi Sugiyama, Rita Singh, Bhiksha Raj |
| 2024 | ICML | Completing Visual Objects via Bridging Generation and Segmentation. | Xiang Li, Yinpeng Chen, Chung-Ching Lin, Hao Chen, Kai Hu, Rita Singh, Bhiksha Raj, Lijuan Wang, Zicheng Liu |
| 2024 | ICPR | Fashion Image Retrieval with Occlusion. | Jimin Sohn, Haeji Jung, Zhiwen Yan, Vibha Masti, Xiang Li, Bhiksha Raj |
| 2024 | Interspeech | SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios. | Hazim T. Bukhari, Soham Deshmukh, Hira Dhamyal, Bhiksha Raj, Rita Singh |
| 2024 | Interspeech | PAM: Prompting Audio-Language Models for Audio Quality Assessment. | Soham Deshmukh, Dareen Alharthi, Benjamin Elizalde, Hannes Gamper, Mahmoud Al Ismail, Rita Singh, Bhiksha Raj, Huaming Wang |
| 2024 | Interspeech | Domain Adaptation for Contrastive Audio-Language Models. | Soham Deshmukh, Rita Singh, Bhiksha Raj |
| 2024 | Interspeech | DeWinder: Single-Channel Wind Noise Reduction using Ultrasound Sensing. | Kuang Yuan, Shuo Han, Swarun Kumar, Bhiksha Raj |
| 2024 | NAACL | AutoPRM: Automating Procedural Supervision for Multi-Step Reasoning via Controllable Question Decomposition. | Zhaorun Chen, Zhuokai Zhao, Zhihong Zhu, Ruiqi Zhang, Xiang Li, Bhiksha Raj, Huaxiu Yao |
| 2024 | NAACL | R-BASS : Relevance-aided Block-wise Adaptation for Speech Summarization. | Roshan Sharma, Ruchira Sharma, Hira Dhamyal, Rita Singh, Bhiksha Raj |
| 2023 | AAAI | Panoramic Video Salient Object Detection with Ambisonic Audio Guidance. | Xiang Li, Haoyuan Cao, Shijie Zhao, Junlin Li, Li Zhang, Bhiksha Raj |
| 2023 | AAAI | VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning. | Kashu Yamazaki, Khoa Vo, Quang Sang Truong, Bhiksha Raj, Ngan Le |
| 2023 | ASRU | Espnet-Summ: Introducing a Novel Large Dataset, Toolkit, and a Cross-Corpora Evaluation of Speech Summarization Systems. | Roshan S. Sharma, William Chen, Takatomo Kano, Ruchira Sharma, Siddhant Arora, Shinji Watanabe, Atsunori Ogawa, Marc Delcroix, Rita Singh, Bhiksha Raj |
| 2023 | CVPR | FREDOM: Fairness Domain Adaptation Approach to Semantic Scene Understanding. | Thanh-Dat Truong, Ngan Le, Bhiksha Raj, Jackson David Cothren, Khoa Luu |
| 2023 | EMNLP | Token Prediction as Implicit Classification to Identify LLM-Generated Text. | Yutian Chen, Hao Kang, Vivian Zhai, Liangze Li, Rita Singh, Bhiksha Raj |
| 2023 | EMNLP | Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and Text. | Xiang Li, Jinglu Wang, Xiaohao Xu, Muqiao Yang, Fan Yang, Yizhou Zhao, Rita Singh, Bhiksha Raj |
| 2023 | ICASSP | An Approach to Ontological Learning from Weak Labels. | Ankit Shah, Larry Tang, Po Hao Chou, Yi Yu Zheng, Ziqian Ge, Bhiksha Raj |
| 2023 | ICASSP | Privacy-Preserving Automatic Speaker Diarization. | Francisco Teixeira, Alberto Abad, Bhiksha Raj, Isabel Trancoso |
| 2023 | ICASSP | Paaploss: A Phonetic-Aligned Acoustic Parameter Loss for Speech Enhancement. | Muqiao Yang, Joseph Konan, David Bick, Yunyang Zeng, Shuo Han, Anurag Kumar, Shinji Watanabe, Bhiksha Raj |
| 2023 | ICASSP | TAPLoss: A Temporal Acoustic Parameter Loss for Speech Enhancement. | Yunyang Zeng, Joseph Konan, Shuo Han, David Bick, Muqiao Yang, Anurag Kumar, Shinji Watanabe, Bhiksha Raj |
| 2023 | ICCV | Robust Referring Video Object Segmentation with Cyclic Structural Consensus. | Xiang Li, Jinglu Wang, Xiaohao Xu, Xiao Li, Bhiksha Raj, Yan Lu |
| 2023 | ICCV | Pairwise Similarity Learning is SimPLE. | Yandong Wen, Weiyang Liu, Yao Feng, Bhiksha Raj, Rita Singh, Adrian Weller, Michael J. Black, Bernhard Schlkopf |
| 2023 | ICLR | SoftMatch: Addressing the Quantity-Quality Tradeoff in Semi-supervised Learning. | Hao Chen, Ran Tao, Yue Fan, Yidong Wang, Jindong Wang, Bernt Schiele, Xing Xie, Bhiksha Raj, Marios Savvides |
| 2023 | ICLR | FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning. | Yidong Wang, Hao Chen, Qiang Heng, Wenxin Hou, Yue Fan, Zhen Wu, Jindong Wang, Marios Savvides, Takahiro Shinozaki, Bhiksha Raj, Bernt Schiele, Xing Xie |
| 2023 | ICML | How Many Perturbations Break This Model? Evaluating Robustness Beyond Adversarial Accuracy. | Raphal Olivier, Bhiksha Raj |
| 2023 | Interspeech | BASS: Block-wise Adaptation for Speech Summarization. | Roshan Sharma, Siddhant Arora, Kenneth Zheng, Shinji Watanabe, Rita Singh, Bhiksha Raj |
| 2023 | Interspeech | There is more than one kind of robustness: Fooling Whisper with adversarial examples. | Raphal Olivier, Bhiksha Raj |
| 2023 | Interspeech | The Hidden Dance of Phonemes and Visage: Unveiling the Enigmatic Link between Phonemes and Facial Features. | Liao Qu, Xianwei Zou, Xiang Li, Yandong Wen, Rita Singh, Bhiksha Raj |
| 2022 | ACSSC | Cross-utterance context for multimodal video transcription. | Roshan Sharma, Bhiksha Raj |
| 2022 | ICLR | SphereFace2: Binary Classification is All You Need for Deep Face Recognition. | Yandong Wen, Weiyang Liu, Adrian Weller, Bhiksha Raj, Rita Singh |
| 2022 | Interspeech | Positional Encoding for Capturing Modality Specific Cadence for Emotion Detection. | Hira Dhamyal, Bhiksha Raj, Rita Singh |
| 2022 | Interspeech | Recent improvements of ASR models in the face of adversarial attacks. | Raphal Olivier, Bhiksha Raj |
| 2022 | Interspeech | Towards End-to-End Private Automatic Speaker Recognition. | Francisco Teixeira, Alberto Abad, Bhiksha Raj, Isabel Trancoso |
| 2022 | Interspeech | Improving Speech Enhancement through Fine-Grained Speech Characteristics. | Muqiao Yang, Joseph Konan, David Bick, Anurag Kumar, Shinji Watanabe, Bhiksha Raj |
| 2021 | BMVC | Point3D: tracking actions as moving points with 3D CNNs. | Shentong Mo, Jingfei Xia, Xiaoqing Tan, Bhiksha Raj |
| 2021 | EMNLP | Sequential Randomized Smoothing for Adversarially Robust Speech Recognition. | Raphal Olivier, Bhiksha Raj |
| 2021 | ICASSP | The in-the-Wild Speech Medical Corpus. | Maria Joana Correia, Francisco Teixeira, Catarina Botelho, Isabel Trancoso, Bhiksha Raj |
| 2021 | ICASSP | High-Frequency Adversarial Defense for Speech and Audio. | Raphal Olivier, Bhiksha Raj, Muhammad A. Shah |
| 2021 | ICASSP | Towards Adversarial Robustness Via Compact Feature Representations. | Muhammad A. Shah, Raphal Olivier, Bhiksha Raj |
| 2021 | ICASSP | FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial Disturbances. | Ali Shahin Shamsabadi, Francisco Seplveda Teixeira, Alberto Abad, Bhiksha Raj, Andrea Cavallaro, Isabel Trancoso |
| 2021 | ICCV | Contrast and Order Representations for Video Self-supervised Learning. | Kai Hu, Jie Shao, Yuan Liu, Bhiksha Raj, Marios Savvides, Zhiqiang Shen |
| 2021 | ICCV | The Right to Talk: An Audio-Visual Transformer Approach. | Thanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham, Bhiksha Raj, Ngan Le, Khoa Luu |
| 2021 | ICCV | Self-Supervised 3D Face Reconstruction via Conditional Estimation. | Yandong Wen, Weiyang Liu, Bhiksha Raj, Rita Singh |
| 2021 | Interspeech | Improving Weakly Supervised Sound Event Detection with Self-Supervised Auxiliary Tasks. | Soham Deshmukh, Bhiksha Raj, Rita Singh |
| 2021 | Interspeech | Masked Proxy Loss for Text-Independent Speaker Verification. | Jiachen Lian, Aiswarya Vinod Kumar, Hira Dhamyal, Bhiksha Raj, Rita Singh |
| 2020 | ICASSP | Deriving Compact Feature Representations Via Annealed Contraction. | Muhammad Ahmed Shah, Bhiksha Raj |
| 2020 | ICPR | Exploiting Non-Linear Redundancy for Neural Model Compression. | Muhammad Ahmed Shah, Raphal Olivier, Bhiksha Raj |
| 2020 | ICPR | Optimal Strategies For Comparing Covariates To Solve Matching Problems. | Muhammad A. Shah, Raphal Olivier, Bhiksha Raj |
| 2020 | ICPR | Hierarchical Routing Mixture of Experts. | Wenbo Zhao, Yang Gao, Shahan Ali Memon, Bhiksha Raj, Rita Singh |
| 2020 | Interspeech | The Phonetic Bases of Vocal Expressed Emotion: Natural versus Acted. | Hira Dhamyal, Shahan Ali Memon, Bhiksha Raj, Rita Singh |
| 2020 | Interspeech | Hide and Speak: Towards Deep Neural Networks for Speech Steganography. | Felix Kreuk, Yossi Adi, Bhiksha Raj, Rita Singh, Joseph Keshet |
| 2020 | ISVC | Controlled AutoEncoders to Generate Faces from Voices. | Hao Liang, Lulan Yu, Guikang Xu, Bhiksha Raj, Rita Singh |
| 2020 | LREC | Automatic In-the-wild Dataset Annotation with Deep Generalized Multiple Instance Learning. | Maria Joana Correia, Isabel Trancoso, Bhiksha Raj |
| 2020 | MASS | Sherlock: A Crowd-sourced System For Automatic Tagging Of Indoor Floor Plans. | Muhammad Ahmed Shah, Khaled A. Harras, Bhiksha Raj |
| 2019 | ASRU | In-the-Wild End-to-End Detection of Speech Affecting Diseases. | M. Joana Correia, Isabel Trancoso, Bhiksha Raj |
| 2019 | ASRU | Optimizing Neural Network Embeddings Using a Pair-Wise Loss for Text-Independent Speaker Verification. | Hira Dhamyal, Tianyan Zhou, Bhiksha Raj, Rita Singh |
| 2019 | ICASSP | Cross Modal Audio Search and Retrieval with Joint Embeddings Based on Text and Audio. | Benjamin Elizalde, Shuayb Zarar, Bhiksha Raj |
| 2019 | ICASSP | Time Signal Classification Using Random Convolutional Features. | Abelino Jimnez, Bhiksha Raj |
| 2019 | ICASSP | Human Behaviour Recognition Using Wifi Channel State Information. | Daanish Ali Khan, Saquib Razak, Bhiksha Raj, Rita Singh |
| 2019 | ICLR | Disjoint Mapping Network for Cross-modal Matching of Voices and Faces. | Yandong Wen, Mahmoud Al Ismail, Weiyang Liu, Bhiksha Raj, Rita Singh |
| 2019 | IJCAI | Learning Sound Events from Webly Labeled Data. | Anurag Kumar, Ankit Shah, Alexander G. Hauptmann, Bhiksha Raj |
| 2019 | IJCNN | Neural Regression Trees. | Shahan Ali Memon, Wenbo Zhao, Bhiksha Raj, Rita Singh |
| 2018 | ICASSP | Framework for Evaluation of Sound Event Detection in Web Videos. | Rohan Badlani, Ankit Shah, Benjamin Elizalde, Anurag Kumar, Bhiksha Raj |
| 2018 | ICASSP | Voice Impersonation Using Generative Adversarial Networks. | Yang Gao, Rita Singh, Bhiksha Raj |
| 2018 | ICASSP | Acoustic Scene Classification Using Discrete Random Hashing for Laplacian Kernel Machines. | Abelino Jimenez, Benjamin Elizalde, Bhiksha Raj |
| 2018 | ICASSP | Content-Based Representations of Audio Using Siamese Neural Networks. | Pranay Manocha, Rohan Badlani, Anurag Kumar, Ankit Shah, Benjamin Elizalde, Bhiksha Raj |
| 2018 | ICASSP | A Corrective Learning Approach for Text-Independent Speaker Verification. | Yandong Wen, Tianyan Zhou, Rita Singh, Bhiksha Raj |
| 2018 | ICMLA | Interactive Evaluation of Classifiers Under Limited Resources. | Sabit Hassan, Shaden Shaar, Bhiksha Raj, Saquib Razak |
| 2018 | Interspeech | Mining Multimodal Repositories for Speech Affecting Diseases. | M. Joana Correia, Bhiksha Raj, Isabel Trancoso, Francisco Teixeira |
| 2018 | PAKDD | Classifier Risk Estimation Under Limited Labeling Resources. | Anurag Kumar, Bhiksha Raj |
| 2017 | CVPR | SphereFace: Deep Hypersphere Embedding for Face Recognition. | Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, Le Song |
| 2017 | ICASSP | Discovering sound concepts and acoustic relations in text. | Anurag Kumar, Bhiksha Raj, Ndapandula Nakashole |
| 2017 | ICASSP | Privacy preserving Distance computation using somewhat-trusted third parties. | Abelino Jimenez, Bhiksha Raj |
| 2017 | ICASSP | Supervised monaural source separation based on autoencoders. | Keiichi Osako, Yuki Mitsufuji, Rita Singh, Bhiksha Raj |
| 2017 | IJCNN | Audio event and scene recognition: A unified approach using strongly and weakly labeled data. | Bhiksha Raj, Anurag Kumar |
| 2017 | Interspeech | Audio Content Based Geotagging in Multimedia. | Anurag Kumar, Benjamin Elizalde, Bhiksha Raj |
| 2017 | Interspeech | Hidden Markov Model Variational Autoencoder for Acoustic Unit Discovery. | Janek Ebbers, Jahn Heymann, Lukas Drude, Thomas Glarner, Reinhold Haeb-Umbach, Bhiksha Raj |
| 2016 | CVPR | The Best of BothWorlds: Combining Data-Independent and Data-Driven Approaches for Action Recognition. | Zhen-Zhong Lan, Shoou-I Yu, Dezhong Yao, Ming Lin, Bhiksha Raj, Alexander G. Hauptmann |
| 2016 | ICASSP | The relationship of voice onset time and Voice Offset Time to physical age. | Rita Singh, Joseph Keshet, Deniz Genaga, Bhiksha Raj |
| 2016 | Interspeech | On the Appropriateness of Complex-Valued Neural Networks for Speech Enhancement. | Lukas Drude, Bhiksha Raj, Reinhold Haeb-Umbach |
| 2016 | ICTD | Viral Spread via Entertainment and Voice-Messaging Among Telephone Users in India. | Agha Ali Raza, Rajat Kulshreshtha, Spandana Gella, Sean Blagsvedt, Maya Chandrasekaran, Bhiksha Raj, Roni Rosenfeld |
| 2015 | ACII | Efficient autism spectrum disorder prediction with eye movement: A machine learning framework. | Wenbo Liu, Li Yi, Zhiding Yu, Xiaobing Zou, Bhiksha Raj, Ming Li |
| 2015 | CVPR | Beyond Gaussian Pyramid: Multi-skip Feature Stacking for action recognition. | Zhen-Zhong Lan, Ming Lin, Xuanchong Li, Alexander G. Hauptmann, Bhiksha Raj |
| 2015 | ICASSP | A novel ranking method for multiple classifier systems. | Anurag Kumar, Bhiksha Raj |
| 2015 | ICASSP | Reducing communication overhead in distributed learning by an order of magnitude (almost). | Anders land, Bhiksha Raj |
| 2015 | ICASSP | Privacy-preserving Query-by-Example Speech Search. | Jos Portelo, Alberto Abad, Bhiksha Raj, Isabel Trancoso |
| 2015 | Interspeech | Locality constrained transitive distance clustering on speech data. | Wenbo Liu, Zhiding Yu, Bhiksha Raj, Ming Li |
| 2014 | ICASSP | Iterative Bayesian word segmentation for unsupervised vocabulary discovery from phoneme lattices. | Jahn Heymann, Oliver Walter, Reinhold Haeb-Umbach, Bhiksha Raj |
| 2014 | ICASSP | Active-set newton algorithm for non-negative sparse coding of audio. | Tuomas Virtanen, Bhiksha Raj, Jort F. Gemmeke, Hugo Van hamme |
| 2014 | Interspeech | Post-masking: a hybrid approach to array processing for speech recognition. | Amir R. Moghimi, Bhiksha Raj, Richard M. Stern |
| 2014 | SIGIR | Privacy-Preserving Important Passage Retrieval. | Lus Marujo, Jos Portelo, David Martins de Matos, Joo Paulo Neto, Anatole Gershman, Jaime G. Carbonell, Isabel Trancoso, Bhiksha Raj |
| 2013 | ASRU | Unsupervised word segmentation from noisy input. | Jahn Heymann, Oliver Walter, Reinhold Haeb-Umbach, Bhiksha Raj |
| 2013 | ASRU | A hierarchical system for word discovery exploiting DTW-based initialization. | Oliver Walter, Timo Korthals, Reinhold Haeb-Umbach, Bhiksha Raj |
| 2013 | ICASSP | Unsupervised hierarchical structure induction for deeper semantic analysis of audio. | Sourish Chaudhuri, Bhiksha Raj |
| 2013 | ICASSP | Optimization of the DET curve in speaker verification under noisy conditions. | Leibny Paola Garca-Perera, Bhiksha Raj, Juan Arturo Nolazco-Flores |
| 2013 | ICASSP | Speaker tracking with spherical microphone arrays. | John W. McDonough, Ken'ichi Kumatani, Takayuki Arakawa, Kazumasa Yamamoto, Bhiksha Raj |
| 2013 | Interspeech | Ensemble approach in speaker verification. | Leibny Paola Garca-Perera, Bhiksha Raj, Juan Arturo Nolazco-Flores |
| 2013 | Interspeech | Discriminatively trained dependency language modeling for conversational speech recognition. | Benjamin Lambert, Bhiksha Raj, Rita Singh |
| 2013 | Interspeech | Secure binary embeddings of front-end factor analysis for privacy preserving speaker verification. | Jos Portelo, Alberto Abad, Bhiksha Raj, Isabel Trancoso |
| 2012 | EACL | An Unsupervised Dynamic Bayesian Network Approach to Measuring Speech Style Accommodation. | Mahaveer Jain, John W. McDonough, Gahgene Gweon, Bhiksha Raj, Carolyn Penstein Ros |
| 2012 | ICASSP | Spectrographic seam patterns for discriminative word spotting. | Shubhranshu Barnwal, Kamal Sahni, Rita Singh, Bhiksha Raj |
| 2012 | ICASSP | Audio event detection from acoustic unit occurrence patterns. | Anurag Kumar, Pranay Dighe, Rita Singh, Sourish Chaudhuri, Bhiksha Raj |
| 2012 | ICASSP | Privacy-preserving speaker verification as password matching. | Manas A. Pathak, Bhiksha Raj |
| 2012 | ICASSP | Attacking a privacy preserving music matching algorithm. | Jos Portelo, Bhiksha Raj, Isabel Trancoso |
| 2012 | Interspeech | Structured sparse coding for microphone array location calibration. | Afsaneh Asaei, Bhiksha Raj, Herv Bourlard, Volkan Cevher |
| 2012 | Interspeech | Exploiting Temporal Sequence Structure for Semantic Analysis of Multimedia. | Sourish Chaudhuri, Rita Singh, Bhiksha Raj |
| 2012 | Interspeech | Plagiarism Detection in Polyphonic Music using Monaural Signal Separation. | Soham De, Indradyumna Roy, Tarunima Prabhakar, Kriti Suneja, Sourish Chaudhuri, Rita Singh, Bhiksha Raj |
| 2012 | Interspeech | Microphone Array Post-filter based on Spatially-Correlated Noise Measurements for Distant Speech Recognition. | Ken'ichi Kumatani, Bhiksha Raj, Rita Singh, John W. McDonough |
| 2012 | Interspeech | Language identification using spectro-temporal patch features. | Kamal Sahni, Pranay Dighe, Rita Singh, Bhiksha Raj |
| 2011 | ACSSC | Greedy sparsity-constrained optimization. | Sohail Bahmani, Petros Boufounos, Bhiksha Raj |
| 2011 | ACSSC | An information filter for voice prompt suppression. | John W. McDonough, Wei Chu, Ken'ichi Kumatani, Bhiksha Raj, Jill Fain Lehman |
| 2011 | ASRU | Maximum kurtosis beamforming with a subspace filter for distant speech recognition. | Ken'ichi Kumatani, John W. McDonough, Bhiksha Raj |
| 2011 | ICASSP | An iterative least-squares technique for dereverberation. | Kshitiz Kumar, Bhiksha Raj, Rita Singh, Richard M. Stern |
| 2011 | ICASSP | Gammatone sub-band magnitude-domain dereverberation for ASR. | Kshitiz Kumar, Rita Singh, Bhiksha Raj, Richard M. Stern |
| 2011 | ICASSP | Privacy preserving probabilistic inference with Hidden Markov Models. | Manas A. Pathak, Shantanu Rane, Wei Sun, Bhiksha Raj |
| 2011 | ICASSP | A paired test for recognizer selection with untranscribed data. | Bhiksha Raj, Rita Singh, James Baker |
| 2011 | Interspeech | Unsupervised Learning of Acoustic Unit Descriptors for Audio Content Representation and Classification. | Sourish Chaudhuri, Mark Harvilla, Bhiksha Raj |
| 2011 | Interspeech | A Paradigm for Limited Vocabulary Speech Recognition Based on Redundant Spectro-Temporal Feature Sets. | Sourish Chaudhuri, Bhiksha Raj, Tony Ezzat |
| 2011 | Interspeech | Privacy Preserving Speaker Verification Using Adapted GMMs. | Manas A. Pathak, Bhiksha Raj |
| 2011 | Interspeech | Phoneme-Dependent NMF for Speech Enhancement in Monaural Mixtures. | Bhiksha Raj, Rita Singh, Tuomas Virtanen |
| 2011 | SIGdial | A Comparison of Latent Variable Models For Conversation Analysis. | Sourish Chaudhuri, Bhiksha Raj |
| 2010 | ICASSP | A hybrid physical and statistical dynamic articulatory framework incorporating analysis-by-synthesis for improved phone classification. | Ziad Al Bawab, Bhiksha Raj, Richard M. Stern |
| 2010 | ICASSP | Learning-based auditory encoding for robust speech recognition. | Yu-Hsiang Bosco Chiu, Bhiksha Raj, Richard M. Stern |
| 2010 | ICASSP | Latent-variable decomposition based dereverberation of monaural and multi-channel signals. | Rita Singh, Bhiksha Raj, Paris Smaragdis |
| 2010 | ICASSP | Ultrasonic sensing for robust speech recognition. | Sundararajan Srinivasan, Bhiksha Raj, Tony Ezzat |
| 2010 | ICASSP | Synthesizing speech from Doppler signals. | Arthur R. Toth, Kaustubh Kalgaonkar, Bhiksha Raj, Tony Ezzat |
| 2010 | ICASSP | Spectrogram dimensionality reductionwith independence constraints. | Kevin W. Wilson, Bhiksha Raj |
| 2010 | Interspeech | Creating a linguistic plausibility dataset with non-expert annotators. | Benjamin Lambert, Rita Singh, Bhiksha Raj |
| 2010 | Interspeech | Non-negative matrix factorization based compensation of music for automatic speech recognition. | Bhiksha Raj, Tuomas Virtanen, Sourish Chaudhuri, Rita Singh |
| 2010 | Interspeech | Ungrounded independent non-negative factor analysis. | Bhiksha Raj, Kevin W. Wilson, Alexander Krueger, Reinhold Haeb-Umbach |
| 2010 | Interspeech | The use of sense in unsupervised training of acoustic models for ASR systems. | Rita Singh, Benjamin Lambert, Bhiksha Raj |
| 2009 | ECIR | Word Particles Applied to Information Retrieval. | Evandro B. Gouva, Bhiksha Raj |
| 2009 | ICASSP | A joint decoding algorithm for multiple-example-based addition of words to a pronunciation lexicon. | Dhananjay Bansal, Nishanth Ulhas Nair, Rita Singh, Bhiksha Raj |
| 2009 | ICASSP | One-handed gesture recognition using ultrasonic Doppler sonar. | Kaustubh Kalgaonkar, Bhiksha Raj |
| 2009 | Interspeech | Deriving vocal tract shapes from electromagnetic articulograph data via geometric adaptation and matching. | Ziad Al Bawab, Lorenzo Turicchia, Richard M. Stern, Bhiksha Raj |
| 2009 | Interspeech | Towards fusion of feature extraction and acoustic model training: a top down process for robust speech recognition. | Yu-Hsiang Bosco Chiu, Bhiksha Raj, Richard M. Stern |
| 2009 | Interspeech | Signal separation for robust speech recognition based on phase difference information obtained in the frequency domain. | Chanwoo Kim, Kshitiz Kumar, Bhiksha Raj, Richard M. Stern |
| 2009 | IDA | Probabilistic Factorization of Non-negative Data with Entropic Co-occurrence Constraints. | Paris Smaragdis, Madhusudana V. S. Shashanka, Bhiksha Raj, Gautham J. Mysore |
| 2008 | ICASSP | Analysis-by-synthesis features for speech recognition. | Ziad Al Bawab, Bhiksha Raj, Richard M. Stern |
| 2008 | ICASSP | Ultrasonic Doppler sensor for speaker recognition. | Kaustubh Kalgaonkar, Bhiksha Raj |
| 2008 | ICASSP | Sparse and shift-invariant feature extraction from non-negative data. | Paris Smaragdis, Bhiksha Raj, Madhusudana V. S. Shashanka |
| 2008 | ICASSP | Speech denoising using nonnegative matrix factorization with priors. | Kevin W. Wilson, Bhiksha Raj, Paris Smaragdis, Ajay Divakaran |
| 2008 | Interspeech | Regularized non-negative matrix factorization with temporal dependencies for speech denoising. | Kevin W. Wilson, Bhiksha Raj, Paris Smaragdis |
| 2007 | AVSS | Acoustic Doppler sonar for gait recogination. | Kaustubh Kalgaonkar, Bhiksha Raj |
| 2007 | CVPR | Sensor and Data Systems, Audio-Assisted Cameras and Acoustic Doppler Sensors. | Kaustubh Kalgaonkar, Paris Smaragdis, Bhiksha Raj |
| 2007 | ICASSP | Bandwidth Expansionwith a plya URN Model. | Bhiksha Raj, Rita Singh, Madhusudana V. S. Shashanka, Paris Smaragdis |
| 2007 | ICASSP | Sparse Overcomplete Decomposition for Single Channel Speaker Separation. | Madhusudana V. S. Shashanka, Bhiksha Raj, Paris Smaragdis |
| 2007 | Interspeech | Probabilistic deduction of symbol mappings for extension of lexicons. | Rita Singh, Evandro B. Gouva, Bhiksha Raj |
| 2006 | ICASSP | Latent Dirichlet Decomposition for Single Channel Speaker Separation. | Bhiksha Raj, Madhusudana V. S. Shashanka, Paris Smaragdis |
| 2006 | Interspeech | An integrated approach to improve speech recognition rate for non-native speakers. | Yunbin Deng, Xiaokun Li, Chiman Kwan, Roger Xu, Bhiksha Raj, Richard M. Stern, David Williamson |
| 2005 | ICASSP | A Companding Front End for Noise-Robust Automatic Speech Recognition. | Jethran Guinness, Bhiksha Raj, Bent Schmidt-Nielsen, Lorenzo Turicchia, Rahul Sarpeshkar |
| 2005 | Interact | A Comparison Between Spoken Queries and Menu-Based Interfaces for In-car Digital Music Selection. | Clifton Forlines, Bent Schmidt-Nielsen, Bhiksha Raj, Kent Wittenburg, Peter Wolf |
| 2005 | Interspeech | Bandwidth expansion of narrowband speech using non-negative matrix factorization. | Dhananjay Bansal, Bhiksha Raj, Paris Smaragdis |
| 2005 | Interspeech | Recognizing speech from simultaneous speakers. | Bhiksha Raj, Rita Singh, Paris Smaragdis |
| 2004 | ICASSP | On tracking noise with linear dynamical system models. | Bhiksha Raj, Rita Singh, Richard M. Stern |
| 2004 | Interspeech | A minimum mean squared error estimator for single channel speaker separation. | Aarthi M. Reddy, Bhiksha Raj |
| 2004 | Interspeech | Soft mask estimation for single channel speaker separation. | Aarthi M. Reddy, Bhiksha Raj |
| 2004 | Interspeech | Spokenquery: an alternate approach to chosing items with speech. | Peter Wolf, Joseph Woelfel, Jan C. van Gemert, Bhiksha Raj, David Wong |
| 2004 | NAACL | A Speech-in List-out Approach to Spoken User Interfaces. | Vijay Divi, Clifton Forlines, Jan C. van Gemert, Bhiksha Raj, Bent Schmidt-Nielsen, Kent Wittenburg, Joseph Woelfel, Fang-Fang Zhang |
| 2003 | ICASSP | Multi-channel source separation by factorial HMMs. | Manuel J. Reyes Gomez, Bhiksha Raj, Dan Ellis |
| 2003 | ICASSP | Lossless compression of language model structure and word identifiers. | Bhiksha Raj, Edward W. D. Whittaker |
| 2003 | ICASSP | Tracking noise via dynamical systems with a continuum of states. | Rita Singh, Bhiksha Raj |
| 2003 | Interspeech | Design of the CMU sphinx-4 decoder. | Paul Lamere, Philip Kwok, William Walker, Evandro B. Gouva, Rita Singh, Bhiksha Raj, Peter Wolf |
| 2003 | Interspeech | Classification with free energy at raised temperatures. | Rita Singh, Manfred K. Warmuth, Bhiksha Raj, Paul Lamere |
| 2002 | ICASSP | Speech recognizer-based microphone array processing for robust hands-free speech recognition. | Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
| 2001 | ICASSP | Speech in Noisy Environments: robust automatic segmentation, feature extraction, and hypothesis combination. | Rita Singh, Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
| 2001 | Interspeech | A boosting approach for confidence scoring. | Pedro J. Moreno, Beth Logan, Bhiksha Raj |
| 2001 | Interspeech | Calibration of microphone arrays for improved speech recognition. | Michael L. Seltzer, Bhiksha Raj |
| 2001 | Interspeech | Quantization-based language model compression. | Edward W. D. Whittaker, Bhiksha Raj |
| 2001 | Interspeech | Comparison of width-wise and length-wise language model compression. | Edward W. D. Whittaker, Bhiksha Raj |
| 2000 | ICASSP | Automatic generation of phone sets and lexical transcriptions. | Rita Singh, Bhiksha Raj, Richard M. Stern |
| 2000 | Interspeech | Reconstruction of damaged spectrographic features for robust speech recognition. | Bhiksha Raj, Michael L. Seltzer, Richard M. Stern |
| 2000 | Interspeech | Classifier-based mask estimation for missing feature methods of robust speech recognition. | Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
| 2000 | Interspeech | Structured redefinition of sound units by merging and splitting for improved speech recognition. | Rita Singh, Bhiksha Raj, Richard M. Stern |
| 1999 | ICASSP | Automatic clustering and generation of contextual questions for tied states in hidden Markov models. | Rita Singh, Bhiksha Raj, Richard M. Stern |
| 1999 | Interspeech | Domain adduced state tying for cross-domain acoustic modelling. | Rita Singh, Bhiksha Raj, Richard M. Stern |
| 1998 | Interspeech | Inference of missing spectrographic features for robust speech recognition. | Bhiksha Raj, Rita Singh, Richard M. Stern |
| 1997 | ICASSP | The effects of background music on speech recognition accuracy. | Bhiksha Raj, Vipul N. Parikh, Richard M. Stern |
| 1996 | ICASSP | A vector Taylor series approach for environment-independent speech recognition. | Pedro J. Moreno, Bhiksha Raj, Richard M. Stern |
| 1996 | Interspeech | Cepstral compensation by polynomial approximation for environment-independent speech recognition. | Bhiksha Raj, Evandro Bacci Gouva, Pedro J. Moreno, Richard M. Stern |
| 1995 | ICASSP | Multivariate-Gaussian-based cepstral normalization for robust speech recognition. | Pedro J. Moreno, Bhiksha Raj, Evandro B. Gouva, Richard M. Stern |
| 1995 | Interspeech | A unified approach for robust speech recognition. | Pedro J. Moreno, Bhiksha Raj, Richard M. Stern |