| 2026 | AAAI | DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis. | Yinghao Aaron Li, Xilin Jiang, Fei Tao, Cheng Niu, Kaifeng Xu, Juntong Song, Nima Mesgarani |
| 2025 | ACL | Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence. | Linyang He, Ercong Nie, Helmut Schmid, Hinrich Schtze, Nima Mesgarani, Jonathan Brennan |
| 2025 | ACL | AAD-LLM: Neural Attention-Driven Auditory Scene Understanding. | Xilin Jiang, Sukru Samet Dindar, Vishal Choudhari, Stephan Bickel, Ashesh D. Mehta, Guy M. McKhann II, Daniel Friedman, Adeen Flinker, Nima Mesgarani |
| 2025 | EMNLP | Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations. | Linyang He, Qiaolin Wang, Xilin Jiang, Nima Mesgarani |
| 2025 | ICASSP | Decoding the Unintelligible: Neural Speech Tracking in Low Signal-to-Noise Ratios. | Xiaomin He, Vinay S. Raghavan, Nima Mesgarani |
| 2025 | ICASSP | Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation. | Xilin Jiang, Cong Han, Nima Mesgarani |
| 2025 | ICASSP | Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis. | Xilin Jiang, Yinghao Aaron Li, Adrian Nicolas Florea, Cong Han, Nima Mesgarani |
| 2025 | Interspeech | Neuro2Semantic: A Transfer Learning Framework for Semantic Reconstruction of Continuous Language from Human Intracranial EEG. | Siavash Shams, Richard J. Antonello, Gavin Mischler, Stephan Bickel, Ashesh D. Mehta, Nima Mesgarani |
| 2025 | NAACL | StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion. | Yinghao Aaron Li, Xilin Jiang, Cong Han, Nima Mesgarani |
| 2024 | ICASSP | Exploring Self-supervised Contrastive Learning of Spatial Sound Event Representation. | Xilin Jiang, Cong Han, Yinghao Aaron Li, Nima Mesgarani |
| 2023 | ICASSP | Online Binaural Speech Separation Of Moving Speakers With A Wavesplit Network. | Cong Han, Nima Mesgarani |
| 2023 | ICASSP | Phoneme-Level Bert for Enhanced Prosody of Text-To-Speech with Grapheme Predictions. | Yinghao Aaron Li, Cong Han, Xilin Jiang, Nima Mesgarani |
| 2023 | Interspeech | DeCoR: Defy Knowledge Forgetting by Predicting Earlier Audio Codes. | Xilin Jiang, Yinghao Aaron Li, Nima Mesgarani |
| 2021 | ICASSP | Speaker and Direction Inferred Dual-Channel Speech Separation. | Chenxing Li, Jiaming Xu, Nima Mesgarani, Bo Xu |
| 2021 | ICASSP | Rethinking The Separation Layers In Speech Separation Networks. | Yi Luo, Zhuo Chen, Cong Han, Chenda Li, Tianyan Zhou, Nima Mesgarani |
| 2021 | ICASSP | Ultra-Lightweight Speech Separation Via Group Communication. | Yi Luo, Cong Han, Nima Mesgarani |
| 2021 | Interspeech | Continuous Speech Separation Using Speaker Inventory for Long Recording. | Cong Han, Yi Luo, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, Zhuo Chen |
| 2021 | Interspeech | Binaural Speech Separation of Moving Speakers With Preserved Spatial Cues. | Cong Han, Yi Luo, Nima Mesgarani |
| 2021 | Interspeech | StarGANv2-VC: A Diverse, Unsupervised, Non-Parallel Framework for Natural-Sounding Voice Conversion. | Yinghao Aaron Li, Ali Zare, Nima Mesgarani |
| 2021 | Interspeech | Empirical Analysis of Generalized Iterative Speech Separation Networks. | Yi Luo, Cong Han, Nima Mesgarani |
| 2021 | Interspeech | Implicit Filter-and-Sum Network for End-to-End Multi-Channel Speech Separation. | Yi Luo, Nima Mesgarani |
| 2020 | ICASSP | Real-Time Binaural Speech Separation with Preserved Spatial Cues. | Cong Han, Yi Luo, Nima Mesgarani |
| 2020 | ICASSP | End-to-end Microphone Permutation and Number Invariant Multi-channel Speech Separation. | Yi Luo, Zhuo Chen, Nima Mesgarani, Takuya Yoshioka |
| 2020 | Interspeech | Separating Varying Numbers of Sources with Auxiliary Autoencoding Loss. | Yi Luo, Nima Mesgarani |
| 2019 | ASRU | FaSNet: Low-Latency Adaptive Beamforming for Multi-Microphone Audio Processing. | Yi Luo, Cong Han, Nima Mesgarani, Enea Ceolini, Shih-Chii Liu |
| 2019 | ICASSP | Augmented Time-frequency Mask Estimation in Cluster-based Source Separation Algorithms. | Yi Luo, Nima Mesgarani |
| 2019 | ICASSP | Online Deep Attractor Network for Real-time Single-channel Speech Separation. | Cong Han, Yi Luo, Nima Mesgarani |
| 2018 | ICASSP | TaSNet: Time-Domain Audio Separation Network for Real-Time, Single-Channel Speech Separation. | Yi Luo, Nima Mesgarani |
| 2018 | ICASSP | Lip2Audspec: Speech Reconstruction from Silent Lip Movements Video. | Hassan Akbari, Himani Arora, Liangliang Cao, Nima Mesgarani |
| 2018 | Interspeech | Real-time Single-channel Dereverberation and Separation with Time-domain Audio Separation Network. | Yi Luo, Nima Mesgarani |
| 2018 | Interspeech | Music Source Activity Detection and Separation Using Deep Attractor Network. | Rajath Kumar, Yi Luo, Nima Mesgarani |
| 2018 | Interspeech | Speech Processing in the Human Brain Meets Deep Learning. | Nima Mesgarani |
| 2017 | ICASSP | Deep attractor network for single-microphone speaker separation. | Zhuo Chen, Yi Luo, Nima Mesgarani |
| 2017 | ICASSP | NAPLib: An open source toolbox for real-time and offline Neural Acoustic Processing. | Bahar Khalighinejad, Tasha Nagamine, Ashesh D. Mehta, Nima Mesgarani |
| 2017 | ICASSP | Deep clustering and conventional networks for music separation: Stronger together. | Yi Luo, Zhuo Chen, John R. Hershey, Jonathan Le Roux, Nima Mesgarani |
| 2017 | ICML | Understanding the Representation and Computation of Multilayer Perceptrons: A Case Study in Speech Recognition. | Tasha Nagamine, Nima Mesgarani |
| 2016 | CogSci | Analyzing distributional learning of phonemic categories in unsupervised deep neural networks. | Okko Rsnen, Tasha Nagamine, Nima Mesgarani |
| 2016 | ICASSP | Synaptic depression in deep neural networks for speech processing. | Wenhao Zhang, Hanyu Li, Minda Yang, Nima Mesgarani |
| 2016 | Interspeech | Adaptation of Neural Networks Constrained by Prior Statistics of Node Co-Activations. | Tasha Nagamine, Zhuo Chen, Nima Mesgarani |
| 2016 | Interspeech | On the Role of Nonlinear Transformations in Deep Neural Network Acoustic Models. | Tasha Nagamine, Michael L. Seltzer, Nima Mesgarani |
| 2015 | Interspeech | Exploring how deep neural networks form phonemic categories. | Tasha Nagamine, Michael L. Seltzer, Nima Mesgarani |
| 2015 | Interspeech | Speech reconstruction from human auditory cortex with deep neural networks. | Minda Yang, Sameer A. Sheth, Catherine A. Schevon, Guy M. McKhann II, Nima Mesgarani |
| 2014 | Interspeech | Principal components of auditory spectro-temporal receptive fields. | Nagaraj Mahajan, Nima Mesgarani, Hynek Hermansky |
| 2013 | ICASSP | Developing a speaker identification system for the DARPA RATS project. | Oldrich Plchot, Spyros Matsoukas, Pavel Matejka, Najim Dehak, Jeff Z. Ma, Sandro Cumani, Ondrej Glembek, Hynek Hermansky, Sri Harish Reddy Mallidi, Nima Mesgarani, Richard M. Schwartz, Mehdi Soufifar, Zheng-Hua Tan, Samuel Thomas, Bing Zhang, Xinhui Zhou |
| 2012 | ICASSP | The UMD-JHU 2011 speaker recognition system. | Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami |
| 2012 | Interspeech | Speech and speaker separation in human auditory cortex. | Nima Mesgarani, Edward Chang |
| 2012 | Interspeech | Developing a Speech Activity Detection System for the DARPA RATS Program. | Tim Ng, Bing Zhang, Long Nguyen, Spyros Matsoukas, Xinhui Zhou, Nima Mesgarani, Karel Vesel, Pavel Matejka |
| 2012 | Interspeech | Acoustic and Data-driven Features for Robust Speech Activity Detection. | Samuel Thomas, Sri Harish Reddy Mallidi, Thomas Janu, Hynek Hermansky, Nima Mesgarani, Xinhui Zhou, Shihab A. Shamma, Tim Ng, Bing Zhang, Long Nguyen, Spyros Matsoukas |
| 2012 | Interspeech | Automatic intelligibility assessment of pathologic speech in head and neck cancer based on auditory-inspired spectro-temporal modulations. | Xinhui Zhou, Daniel Garcia-Romero, Nima Mesgarani, Maureen L. Stone, Carol Y. Espy-Wilson, Shihab A. Shamma |
| 2011 | ICASSP | Speech processing with a cortical representation of audio. | Nima Mesgarani, Shihab A. Shamma |
| 2011 | Interspeech | Adaptive Stream Fusion in Multistream Recognition of Speech. | Nima Mesgarani, Samuel Thomas, Hynek Hermansky |
| 2010 | ICASSP | Nonlinear filtering of spectrotemporal modulations in speech enhancement. | Majid Mirbagheri, Nima Mesgarani, Shihab A. Shamma |
| 2010 | Interspeech | A multistream multiresolution framework for phoneme recognition. | Nima Mesgarani, Samuel Thomas, Hynek Hermansky |
| 2010 | Interspeech | A phoneme recognition framework based on auditory spectro-temporal receptive fields. | Samuel Thomas, Kailash Patil, Sriram Ganapathy, Nima Mesgarani, Hynek Hermansky |
| 2010 | ISCAS | The use of spike-based representations for hardware audition systems. | Shih-Chii Liu, Nima Mesgarani, John G. Harris, Hynek Hermansky |
| 2009 | Interspeech | Discriminant spectrotemporal features for phoneme recognition. | Nima Mesgarani, Garimella S. V. S. Sivaram, Sridhar Krishna Nemala, Mounya Elhilali, Hynek Hermansky |
| 2007 | ICASSP | Representation of Phonemes in Primary Auditory Cortex: How the Brain Analyzes Speech. | Nima Mesgarani, Stephen V. David, Shihab A. Shamma |
| 2006 | Interspeech | Discriminating speech and non-speech with regularized least squares. | Ryan Rifkin, Nima Mesgarani |
| 2005 | ICASSP | Speech Enhancement Based on Filtering the Spectrotemporal Modulations. | Nima Mesgarani, Shihab A. Shamma |
| 2004 | ICASSP | Speech discrimination based on multiscale spectro-temporal modulations. | Nima Mesgarani, Shihab A. Shamma, Malcolm Slaney |