| 2026 | LREC | Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems. | Maliha Jahan, Thomas Thebaud, Zsuzsanna Fagyal, Jess Villalba, Mark Hasegawa-Johnson, Laureano Moro-Velzquez, Najim Dehak |
| 2025 | ASRU | The JHU-MIT System for NIST SRE24: Post-Evaluation Analysis. | Jess Villalba, Jonas Borgstrom, Prabhav Singh, Leibny Paola Garca, Pedro A. Torres-Carrasquillo, Najim Dehak |
| 2025 | EMNLP | Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation. | Yen-Ju Lu, Thomas Thebaud, Laureano Moro-Velzquez, Najim Dehak, Jess Villalba |
| 2025 | ICASSP | Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and More. | Maliha Jahan, Priyam Mazumdar, Thomas Thebaud, Mark Hasegawa-Johnson, Jess Villalba, Najim Dehak, Laureano Moro-Velzquez |
| 2025 | ICASSP | Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings. | Sarah Laouedj, Yuzhe Wang, Jess Villalba, Thomas Thebaud, Laureano Moro-Velzquez, Najim Dehak |
| 2025 | Interspeech | FaiST: A Benchmark Dataset for Fairness in Speech Technology. | Maliha Jahan, Yinglun Sun, Priyam Mazumdar, Zsuzsanna Fagyal, Thomas Thebaud, Jess Villalba, Mark Hasegawa-Johnson, Najim Dehak, Laureano Moro-Velzquez |
| 2025 | Interspeech | EmoJudge: LLM Based Post-Hoc Refinement for Multimodal Speech Emotion Recognition. | Prabhav Singh, Jess Villalba |
| 2025 | Interspeech | Count Your Speakers! Multitask Learning for Multimodal Speaker Diarization. | Prabhav Singh, Jess Villalba, Najim Dehak |
| 2025 | Interspeech | Multimodal Emotion Diarization: Frame-Wise Integration of Text and Audio Representations. | Ziv Tamir, Thomas Thebaud, Jess Villalba, Najim Dehak, Oren Kurland |
| 2024 | ICMI | Multimodal Emotion Recognition Harnessing the Complementarity of Speech, Language, and Vision. | Thomas Thebaud, Anna Favaro, Yaohan Guan, Yuchen Yang, Prabhav Singh, Jess Villalba, Laureano Moro-Velzquez, Najim Dehak |
| 2024 | Interspeech | Noise-robust Speech Separation with Fast Generative Correction. | Helin Wang, Jess Villalba, Laureano Moro-Velzquez, Jiarui Hai, Thomas Thebaud, Najim Dehak |
| 2024 | Interspeech | Exploring the Complementary Nature of Speech and Eye Movements for Profiling Neurological Disorders. | Yuzhe Wang, Anna Favaro, Thomas Thebaud, Jess Villalba, Najim Dehak, Laureano Moro-Velzquez |
| 2023 | ASRU | Model-Based Fairness Metric for Speaker Verification. | Maliha Jahan, Laureano Moro-Velzquez, Thomas Thebaud, Najim Dehak, Jess Villalba |
| 2023 | ASRU | Joint Energy-Based Model for Robust Speech Classification System Against Dirty-Label Backdoor Poisoning Attacks. | Martin Sustek, Sonal Joshi, Henry Li, Thomas Thebaud, Jess Villalba, Sanjeev Khudanpur, Najim Dehak |
| 2023 | ASRU | Clustering Unsupervised Representations as Defense Against Poisoning Attacks on Speech Commands Classification System. | Thomas Thebaud, Sonal Joshi, Henry Li, Martin Sustek, Jess Villalba, Sanjeev Khudanpur, Najim Dehak |
| 2023 | Interspeech | Advances in Language Recognition in Low Resource African Languages: The JHU-MIT Submission for NIST LRE22. | Jess Villalba, Jonas Borgstrom, Maliha Jahan, Saurabh Kataria, Leibny Paola Garca, Pedro A. Torres-Carrasquillo, Najim Dehak |
| 2023 | Interspeech | Segmental SpeechCLIP: Utilizing Pretrained Image-text Models for Audio-Visual Learning. | Saurabhchand Bhati, Jess Villalba, Laureano Moro-Velzquez, Thomas Thebaud, Najim Dehak |
| 2023 | Interspeech | Do Phonatory Features Display Robustness to Characterize Parkinsonian Speech Across Corpora? | Anna Favaro, Tianyu Cao, Thomas Thebaud, Jess Villalba, Ankur A. Butala, Najim Dehak, Laureano Moro-Velzquez |
| 2023 | Interspeech | Self-FiLM: Conditioning GANs with self-supervised representations for bandwidth extension based speaker recognition. | Saurabh Kataria, Jess Villalba, Laureano Moro-Velzquez, Thomas Thebaud, Najim Dehak |
| 2023 | Interspeech | DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model. | Helin Wang, Thomas Thebaud, Jess Villalba, Myra Sydnor, Becky Lammers, Najim Dehak, Laureano Moro-Velzquez |
| 2022 | Interspeech | Non-contrastive self-supervised learning of utterance-level speech representations. | Jaejin Cho, Raghavendra Pappagari, Piotr Zelasko, Laureano Moro-Velzquez, Jess Villalba, Najim Dehak |
| 2022 | Interspeech | Defense against Adversarial Attacks on Hybrid Speech Recognition System using Adversarial Fine-tuning with Denoiser. | Sonal Joshi, Saurabh Kataria, Yiwen Shao, Piotr Zelasko, Jess Villalba, Sanjeev Khudanpur, Najim Dehak |
| 2022 | Interspeech | AdvEst: Adversarial Perturbation Estimation to Classify and Detect Adversarial Attacks against Speaker Identification. | Sonal Joshi, Saurabh Kataria, Jess Villalba, Najim Dehak |
| 2022 | Interspeech | Joint domain adaptation and speech bandwidth extension using time-domain GANs for speaker verification. | Saurabh Kataria, Jess Villalba, Laureano Moro-Velzquez, Najim Dehak |
| 2022 | Interspeech | End-to-End Neural Speaker Diarization with an Iterative Refinement of Non-Autoregressive Attention-based Attractors. | Magdalena Rybicka, Jess Villalba, Najim Dehak, Konrad Kowalczyk |
| 2022 | Interspeech | Chunking Defense for Adversarial Attacks on ASR. | Yiwen Shao, Jess Villalba, Sonal Joshi, Saurabh Kataria, Sanjeev Khudanpur, Najim Dehak |
| 2021 | ASRU | Beyond Isolated Utterances: Conversational Emotion Recognition. | Raghavendra Pappagari, Piotr Zelasko, Jess Villalba, Laureano Moro-Velzquez, Najim Dehak |
| 2021 | ICASSP | Focus on the Present: A Regularization Method for the ASR Source-Target Attention Layer. | Nanxin Chen, Piotr Zelasko, Jess Villalba, Najim Dehak |
| 2021 | ICASSP | Improving Reconstruction Loss Based Speaker Embedding in Unsupervised and Semi-Supervised Scenarios. | Jaejin Cho, Piotr Zelasko, Jess Villalba, Najim Dehak |
| 2021 | ICASSP | Perceptual Loss Based Speech Denoising with an Ensemble of Audio Pattern Recognition and Self-Supervised Models. | Saurabh Kataria, Jess Villalba, Najim Dehak |
| 2021 | ICASSP | CopyPaste: An Augmentation Method for Speech Emotion Recognition. | Raghavendra Pappagari, Jess Villalba, Piotr Zelasko, Laureano Moro-Velzquez, Najim Dehak |
| 2021 | Interspeech | Segmental Contrastive Predictive Coding for Unsupervised Word Segmentation. | Saurabhchand Bhati, Jess Villalba, Piotr Zelasko, Laureano Moro-Velzquez, Najim Dehak |
| 2021 | Interspeech | Align-Denoise: Single-Pass Non-Autoregressive Speech Recognition. | Nanxin Chen, Piotr Zelasko, Laureano Moro-Velzquez, Jess Villalba, Najim Dehak |
| 2021 | Interspeech | Deep Feature CycleGANs: Speaker Identity Preserving Non-Parallel Microphone-Telephone Domain Adaptation for Speaker Verification. | Saurabh Kataria, Jess Villalba, Piotr Zelasko, Laureano Moro-Velzquez, Najim Dehak |
| 2021 | Interspeech | Automatic Detection and Assessment of Alzheimer Disease Using Speech and Language Technologies in Low-Resource Scenarios. | Raghavendra Pappagari, Jaejin Cho, Sonal Joshi, Laureano Moro-Velzquez, Piotr Zelasko, Jess Villalba, Najim Dehak |
| 2021 | Interspeech | Spine2Net: SpineNet with Res2Net and Time-Squeeze-and-Excitation Blocks for Speaker Recognition. | Magdalena Rybicka, Jess Villalba, Piotr Zelasko, Najim Dehak, Konrad Kowalczyk |
| 2021 | Interspeech | Representation Learning to Classify and Detect Adversarial Attacks Against Speaker and Speech Recognition Systems. | Jess Villalba, Sonal Joshi, Piotr Zelasko, Najim Dehak |
| 2020 | ICASSP | Feature Enhancement with Deep Feature Losses for Speaker Verification. | Saurabh Kataria, Phani Sankar Nidadavolu, Jess Villalba, Nanxin Chen, L. Paola Garca-Perera, Najim Dehak |
| 2020 | ICASSP | Using X-Vectors to Automatically Detect Parkinson's Disease from Speech. | Laureano Moro-Velzquez, Jess Villalba, Najim Dehak |
| 2020 | ICASSP | Unsupervised Feature Enhancement for Speaker Verification. | Phani Sankar Nidadavolu, Saurabh Kataria, Jess Villalba, L. Paola Garca-Perera, Najim Dehak |
| 2020 | ICASSP | X-Vectors Meet Emotions: A Study On Dependencies Between Emotion and Speaker Recognition. | Raghavendra Pappagari, Tianzi Wang, Jess Villalba, Nanxin Chen, Najim Dehak |
| 2020 | Interspeech | Self-Expressing Autoencoders for Unsupervised Spoken Term Discovery. | Saurabhchand Bhati, Jess Villalba, Piotr Zelasko, Najim Dehak |
| 2020 | Interspeech | Learning Speaker Embedding from Text-to-Speech. | Jaejin Cho, Piotr Zelasko, Jess Villalba, Shinji Watanabe, Najim Dehak |
| 2020 | Interspeech | x-Vectors Meet Adversarial Attacks: Benchmarking Adversarial Robustness in Speaker Verification. | Jess Villalba, Yuekai Zhang, Najim Dehak |
| 2020 | Interspeech | Black-Box Attacks on Spoofing Countermeasures Using Transferability of Adversarial Examples. | Yuekai Zhang, Ziyan Jiang, Jess Villalba, Najim Dehak |
| 2019 | ASRU | Low-Resource Domain Adaptation for Speaker Recognition Using Cycle-Gans. | Phani Sankar Nidadavolu, Saurabh Kataria, Jess Villalba, Najim Dehak |
| 2019 | ASRU | Hierarchical Transformers for Long Document Classification. | Raghavendra Pappagari, Piotr Zelasko, Jess Villalba, Yishay Carmiel, Najim Dehak |
| 2019 | ICASSP | Language Model Integration Based on Memory Control for Sequence to Sequence Speech Recognition. | Jaejin Cho, Shinji Watanabe, Takaaki Hori, Murali Karthick Baskar, Hirofumi Inaguma, Jess Villalba, Najim Dehak |
| 2019 | ICASSP | Investigation on Neural Bandwidth Extension of Telephone Speech for Improved Speaker Recognition. | Phani Sankar Nidadavolu, Vicente Iglesias, Jess Villalba, Najim Dehak |
| 2019 | ICASSP | Cycle-GANs for Domain Adaptation of Acoustic Features for Speaker Recognition. | Phani Sankar Nidadavolu, Jess Villalba, Najim Dehak |
| 2019 | Interspeech | Tied Mixture of Factor Analyzers Layer to Combine Frame Level Representations in Neural Speaker Embeddings. | Nanxin Chen, Jess Villalba, Najim Dehak |
| 2019 | Interspeech | ASSERT: Anti-Spoofing with Squeeze-Excitation and Residual Networks. | Cheng-I Lai, Nanxin Chen, Jess Villalba, Najim Dehak |
| 2019 | Interspeech | The JHU Speaker Recognition System for the VOiCES 2019 Challenge. | David Snyder, Jess Villalba, Nanxin Chen, Daniel Povey, Gregory Sell, Najim Dehak, Sanjeev Khudanpur |
| 2019 | Interspeech | State-of-the-Art Speaker Recognition for Telephone and Video Speech: The JHU-MIT Submission for NIST SRE18. | Jess Villalba, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, Franois Grondin, Rda Dehak, Leibny Paola Garca-Perera, Daniel Povey, Pedro A. Torres-Carrasquillo, Sanjeev Khudanpur, Najim Dehak |
| 2018 | ICASSP | Measuring Uncertainty in Deep Regression Models: The Case of Age Estimation from Speech. | Nanxin Chen, Jess Villalba, Yishay Carmiel, Najim Dehak |
| 2018 | ICASSP | Joint Verification-Identification in end-to-end Multi-Scale CNN Framework for Topic Identification. | Raghavendra Pappagari, Jess Villalba, Najim Dehak |
| 2018 | Interspeech | An Investigation of Non-linear i-vectors for Speaker Verification. | Nanxin Chen, Jess Villalba, Najim Dehak |
| 2018 | Interspeech | Deep Neural Networks for Emotion Recognition Combining Audio and Transcripts. | Jaejin Cho, Raghavendra Pappagari, Purva Kulkarni, Jess Villalba, Yishay Carmiel, Najim Dehak |
| 2018 | Interspeech | Effectiveness of Single-Channel BLSTM Enhancement for Language Identification. | Peter Sibbern Frederiksen, Jess Villalba, Shinji Watanabe, Zheng-Hua Tan, Najim Dehak |
| 2018 | Interspeech | End-to-end Deep Neural Network Age Estimation. | Pegah Ghahremani, Phani Sankar Nidadavolu, Nanxin Chen, Jess Villalba, Daniel Povey, Sanjeev Khudanpur, Najim Dehak |
| 2018 | Interspeech | Investigation on Bandwidth Extension for Speaker Recognition. | Phani Sankar Nidadavolu, Cheng-I Lai, Jess Villalba, Najim Dehak |
| 2018 | Interspeech | Diarization is Hard: Some Experiences and Lessons Learned for the JHU Team in the Inaugural DIHARD Challenge. | Gregory Sell, David Snyder, Alan McCree, Daniel Garcia-Romero, Jess Villalba, Matthew Maciejewski, Vimal Manohar, Najim Dehak, Daniel Povey, Shinji Watanabe, Sanjeev Khudanpur |
| 2017 | Interspeech | Tied Variational Autoencoder Backends for i-Vector Speaker Recognition. | Jess Villalba, Niko Brmmer, Najim Dehak |