| 2026 | LREC | Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems. | Maliha Jahan, Thomas Thebaud, Zsuzsanna Fagyal, Jess Villalba, Mark Hasegawa-Johnson, Laureano Moro-Velzquez, Najim Dehak |
| 2025 | CHI | Speech AI for All: Promoting Accessibility, Fairness, Inclusivity, and Equity. | Shaomei Wu, Kimi Wenzel, Jingjin Li, Qisheng Li, Alisha Pradhan, Raja S. Kushalnagar, Colin Lea, Allison Koenecke, Christian Vogler, Mark Hasegawa-Johnson, Norman Makoto Su, Nan Bernstein Ratner |
| 2025 | ICASSP | Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and More. | Maliha Jahan, Priyam Mazumdar, Thomas Thebaud, Mark Hasegawa-Johnson, Jess Villalba, Najim Dehak, Laureano Moro-Velzquez |
| 2025 | ICASSP | Improved Recognition of the Speech of People with Parkinson's Who Stutter. | Jonghwan Na, Xiuwen Zheng, Bowon Lee, Mark Hasegawa-Johnson |
| 2025 | ICASSP | Cohort-Sensitive Labeling: An Effective Approach for Enhancing ASR Performance. | Jonghwan Na, Mark Hasegawa-Johnson, Bowon Lee |
| 2025 | ICASSP | LIMMITS'25: Multilingual Streaming TTS With Neural Codecs for Indian Languages. | Philipp Olbrich, Hema A. Murthy, Pranaw Kumar, Shinji Watanabe, Sheng Zhao, Mark Hasegawa-Johnson |
| 2025 | ICASSP | Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition. | Satwinder Singh, Qianli Wang, Zihan Zhong, Clarion Mendes, Mark Hasegawa-Johnson, Waleed Abdulla, Seyed Reza Shahamiri |
| 2025 | ICASSP | Dysarthric Speech Conformer: Adaptation for Sequence-to-Sequence Dysarthric Speech Recognition. | Qianli Wang, Zihan Zhong, Satwinder Singh, Clarion Mendes, Mark Hasegawa-Johnson, Waleed Abdulla, Seyed Reza Shahamiri |
| 2025 | Interspeech | The Interspeech 2025 Speech Accessibility Project Challenge. | Xiuwen Zheng, Bornali Phukon, Jonghwan Na, Ed Cutrell, Kyu J. Han, Mark Hasegawa-Johnson, Pan-Pan Jiang, Aadhrik Kuila, Colin Lea, Bob MacDonald, Gautam Varma Mantena, Venkatesh Ravichandran, Leda Sari, Katrin Tomanek, Chang D. Yoo, Chris Zwilling |
| 2025 | Interspeech | SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment. | SooHwan Eom, Mark Hasegawa-Johnson, Chang D. Yoo |
| 2025 | Interspeech | Band-Split Self-supervised Mamba for Infant-centered Audio Analysis. | Xulin Fan, Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain |
| 2025 | Interspeech | FaiST: A Benchmark Dataset for Fairness in Speech Technology. | Maliha Jahan, Yinglun Sun, Priyam Mazumdar, Zsuzsanna Fagyal, Thomas Thebaud, Jess Villalba, Mark Hasegawa-Johnson, Najim Dehak, Laureano Moro-Velzquez |
| 2025 | Interspeech | Aligning ASR Evaluation with Human and LLM Judgments: Intelligibility Metrics Using Phonetic, Semantic, and NLI Approaches. | Bornali Phukon, Xiuwen Zheng, Mark Hasegawa-Johnson |
| 2025 | Interspeech | The Speech Accessibility Project: Best Practices for Collection and Curation of Disordered Speech. | Chris Zwilling, Mark Hasegawa-Johnson, Heather Hodges, Lorraine O. Ramig, Adina Bradshaw, Clarion Mendes, Heejin Kim, Alexandria Barkhimer, Laura Mattie, Meg Dickinson, Shawnise Carter, Marie Moore Channell |
| 2025 | WACV | SyncDiff: Diffusion-Based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved Synchronization. | Xulin Fan, Heting Gao, Ziyi Chen, Peng Chang, Mei Han, Mark Hasegawa-Johnson |
| 2024 | ACL | TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback. | Eunseop Yoon, Hee Suk Yoon, SooHwan Eom, Gunsoo Han, Daniel Wontae Nam, Daejin Jo, Kyoung-Woon On, Mark Hasegawa-Johnson, Sungwoong Kim, Chang Dong Yoo |
| 2024 | CHASE | Sound Tagging in Infant-centric Home Soundscapes. | Mohammad Nur Hossain Khan, Jialu Li, Nancy L. McElwain, Mark Hasegawa-Johnson, Bashima Islam |
| 2024 | COLING | Finding Spoken Identifications: Using GPT-4 Annotation for an Efficient and Fast Dataset Creation Pipeline. | Maliha Jahan, Helin Wang, Thomas Thebaud, Yinglun Sun, Giang Ha Le, Zsuzsanna Fagyal, Odette Scharenborg, Mark Hasegawa-Johnson, Laureano Moro-Velzquez, Najim Dehak |
| 2024 | EMNLP | Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM. | SooHwan Eom, Jay Shim, Gwanhyeong Koo, Haebin Na, Mark Hasegawa-Johnson, Sungwoong Kim, Chang Dong Yoo |
| 2024 | ICASSP | AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition. | SooHwan Eom, Eunseop Yoon, Hee Suk Yoon, Chanwoo Kim, Mark Hasegawa-Johnson, Chang D. Yoo |
| 2024 | ICASSP | G2PU: Grapheme-To-Phoneme Transducer with Speech Units. | Heting Gao, Mark Hasegawa-Johnson, Chang D. Yoo |
| 2024 | ICASSP | Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations. | Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain |
| 2024 | ICASSP | LIMMITS'24: Multi-Speaker, Multi-Lingual Indic TTS with Voice Cloning. | Abhayjeet Singh, Amala Nagireddi, Deekshitha G, Jesuraja Bandekar, Roopa R., Sandhya Badiger, Sathvik Udupa, Prasanta Kumar Ghosh, Hema A. Murthy, Pranaw Kumar, Keiichi Tokuda, Mark Hasegawa-Johnson, Philipp Olbrich |
| 2024 | ICASSP | Unsupervised Speech Recognition with N-skipgram and Positional Unigram Matching. | Liming Wang, Mark Hasegawa-Johnson, Chang D. Yoo |
| 2024 | BSN | InfantMotion2Vec: Unlabeled Data-Driven Infant Pose Estimation Using a Single Chest IMU. | Mohammad Nur Hossain Khan, Nancy L. McElwain, Mark Hasegawa-Johnson, Bashima Islam |
| 2024 | Interspeech | Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis. | Jialu Li, Mark Hasegawa-Johnson, Karrie Karahalios |
| 2024 | Interspeech | Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility. | Xiuwen Zheng, Bornali Phukon, Mark Hasegawa-Johnson |
| 2024 | Interspeech | Visualization for improving foreign language pronunciation. | Charlotte Yoder, Karrie Karahalios, Mark Hasegawa-Johnson, Shreyansh Agrawal |
| 2024 | Interspeech | LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition. | Eunseop Yoon, Hee Suk Yoon, John B. Harvill, Mark Hasegawa-Johnson, Chang D. Yoo |
| 2023 | ACL | A Theory of Unsupervised Speech Recognition. | Liming Wang, Mark Hasegawa-Johnson, Chang Dong Yoo |
| 2023 | ACL | Listen, Decipher and Sign: Toward Unsupervised Speech-to-Sign Language Recognition. | Liming Wang, Junrui Ni, Heting Gao, Jialu Li, Kai Chieh Chang, Xulin Fan, Junkai Wu, Mark Hasegawa-Johnson, Chang Dong Yoo |
| 2023 | ACL | INTapt: Information-Theoretic Adversarial Prompt Tuning for Enhanced Non-Native Speech Recognition. | Eunseop Yoon, Hee Suk Yoon, John B. Harvill, Mark Hasegawa-Johnson, Chang Dong Yoo |
| 2023 | ICASSP | Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-Speech. | Abhayjeet Singh, Amala Nagireddi, Deekshitha G, Jesuraja Bandekar, Roopa R., Sandhya Badiger, Sathvik Udupa, Prasanta Kumar Ghosh, Hema A. Murthy, Heiga Zen, Pranaw Kumar, Kamal Kant, Amol Bole, Bira Chandra Singh, Keiichi Tokuda, Mark Hasegawa-Johnson, Philipp Olbrich |
| 2023 | ICASSP | Dual-Path Cross-Modal Attention for Better Audio-Visual Speech Extraction. | Zhongweiyang Xu, Xulin Fan, Mark Hasegawa-Johnson |
| 2023 | Interspeech | End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions. | Wonjune Kang, Mark Hasegawa-Johnson, Deb Roy |
| 2023 | Interspeech | Towards Robust Family-Infant Audio Analysis Based on Unsupervised Pretraining of Wav2vec 2.0 on Large-Scale Unlabeled Family Audio. | Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain |
| 2023 | Interspeech | Mitigating the Exposure Bias in Sentence-Level Grapheme-to-Phoneme (G2P) Transduction. | Eunseop Yoon, Hee Suk Yoon, Dhananjaya Gowda, SooHwan Eom, Daehyeok Kim, John B. Harvill, Heting Gao, Mark Hasegawa-Johnson, Chanwoo Kim, Chang D. Yoo |
| 2023 | Interspeech | Wav2ToBI: a new approach to automatic ToBI transcription. | Wanyue Zhai, Mark Hasegawa-Johnson |
| 2022 | AAAI | Fast and Efficient MMD-Based Fair PCA via Optimization over Stiefel Manifold. | Junghyun Lee, Gwangsu Kim, Mahbod Olfat, Mark Hasegawa-Johnson, Chang D. Yoo |
| 2022 | ACL | Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech Recognition. | Liming Wang, Siyuan Feng, Mark Hasegawa-Johnson, Chang Dong Yoo |
| 2022 | AISTATS | Equivariance Discovery by Learned Parameter-Sharing. | Raymond A. Yeh, Yuan-Ting Hu, Mark Hasegawa-Johnson, Alexander G. Schwing |
| 2022 | EMNLP | SMSMix: Sense-Maintained Sentence Mixup for Word Sense Disambiguation. | Hee Suk Yoon, Eunseop Yoon, John B. Harvill, Sunjae Yoon, Mark Hasegawa-Johnson, Chang Dong Yoo |
| 2022 | ICASSP | SpeechSplit2.0: Unsupervised Speech Disentanglement for Voice Conversion without Tuning Autoencoder Bottlenecks. | Chak Ho Chan, Kaizhi Qian, Yang Zhang, Mark Hasegawa-Johnson |
| 2022 | ICASSP | Detection of Covid-19 from Joint Time and Frequency Analysis of Speech, Breathing and Cough Audio. | John B. Harvill, Yash R. Wani, Moitreya Chatterjee, Mustafa Alam, David G. Beiser, David Chestek, Mark Hasegawa-Johnson, Narendra Ahuja |
| 2022 | ICML | Forget-free Continual Learning with Winning Subnetworks. | Haeyong Kang, Rusty John Lloyd Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon, Mark Hasegawa-Johnson, Sung Ju Hwang, Chang D. Yoo |
| 2022 | ICML | ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers. | Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni, Cheng-I Lai, David D. Cox, Mark Hasegawa-Johnson, Shiyu Chang |
| 2022 | Interspeech | WavPrompt: Towards Few-Shot Spoken Language Understanding with Frozen Language Models. | Heting Gao, Junrui Ni, Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson |
| 2022 | Interspeech | Frame-Level Stutter Detection. | John B. Harvill, Mark Hasegawa-Johnson, Chang D. Yoo |
| 2022 | Interspeech | Cross-lingual articulatory feature information transfer for speech recognition using recurrent progressive neural networks. | Mahir Morshed, Mark Hasegawa-Johnson |
| 2022 | Interspeech | Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition. | Junrui Ni, Liming Wang, Heting Gao, Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson |
| 2022 | NAACL | Syn2Vec: Synset Colexification Graphs for Lexical Semantic Similarity. | John B. Harvill, Roxana Girju, Mark Hasegawa-Johnson |
| 2021 | ACSSC | A Translation Framework for Visually Grounded Spoken Unit Discovery. | Liming Wang, Mark Hasegawa-Johnson |
| 2021 | ICASSP | How Phonotactics Affect Multilingual and Zero-Shot ASR Performance. | Siyuan Feng, Piotr Zelasko, Laureano Moro-Velzquez, Ali Abavisani, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
| 2021 | ICASSP | Synthesis of New Words for Improved Dysarthric Speech Recognition on an Expanded Vocabulary. | John B. Harvill, Dias Issa, Mark Hasegawa-Johnson, Chang Dong Yoo |
| 2021 | ICASSP | Continuous Cnn For Nonuniform Time Series. | Hui Shi, Yang Zhang, Hao Wu, Shiyu Chang, Kaizhi Qian, Mark Hasegawa-Johnson, Jishen Zhao |
| 2021 | ICASSP | Show and Speak: Directly Synthesize Spoken Description of Images. | Xinsheng Wang, Siyuan Feng, Jihua Zhu, Mark Hasegawa-Johnson, Odette Scharenborg |
| 2021 | ICASSP | Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image Retrieval. | Liming Wang, Xinsheng Wang, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
| 2021 | ICASSP | A Comparison Study on Infant-Parent Voice Diarization. | Junzhe Zhu, Mark Hasegawa-Johnson, Nancy L. McElwain |
| 2021 | ICASSP | Multi-Decoder Dprnn: Source Separation for Variable Number of Speakers. | Junzhe Zhu, Raymond A. Yeh, Mark Hasegawa-Johnson |
| 2021 | ICCV | Interpretable Visual Reasoning via Induced Symbolic Space. | Zhonghao Wang, Kai Wang, Mo Yu, Jinjun Xiong, Wen-Mei Hwu, Mark Hasegawa-Johnson, Humphrey Shi |
| 2021 | ICML | Global Prosody Style Transfer Without Text Transcriptions. | Kaizhi Qian, Yang Zhang, Shiyu Chang, Jinjun Xiong, Chuang Gan, David D. Cox, Mark Hasegawa-Johnson |
| 2021 | Interspeech | Zero-Shot Cross-Lingual Phonetic Recognition with External Language Embedding. | Heting Gao, Junrui Ni, Yang Zhang, Kaizhi Qian, Shiyu Chang, Mark Hasegawa-Johnson |
| 2021 | Interspeech | Classification of COVID-19 from Cough Using Autoregressive Predictive Coding Pretraining and Spectral Data Augmentation. | John B. Harvill, Yash R. Wani, Mark Hasegawa-Johnson, Narendra Ahuja, David G. Beiser, David Chestek |
| 2021 | NAACL | Worldly Wise (WoW) - Cross-Lingual Knowledge Fusion for Fact-based Visual Spoken-Question Answering. | Kiran Ramnath, Leda Sari, Mark Hasegawa-Johnson, Chang D. Yoo |
| 2020 | ICASSP | F0-Consistent Many-To-Many Non-Parallel Voice Conversion Via Conditional Autoencoder. | Kaizhi Qian, Zeyu Jin, Mark Hasegawa-Johnson, Gautham J. Mysore |
| 2020 | ICASSP | Training Spoken Language Understanding Systems with Non-Parallel Speech and Text. | Leda Sari, Samuel Thomas, Mark Hasegawa-Johnson |
| 2020 | ICML | Unsupervised Speech Decomposition via Triple Information Bottleneck. | Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson, David D. Cox |
| 2020 | Interspeech | Automatic Estimation of Intelligibility Measure for Consonants in Speech. | Ali Abavisani, Mark Hasegawa-Johnson |
| 2020 | Interspeech | Evaluating Automatically Generated Phoneme Captions for Images. | Justin van der Hout, Zoltn D'Haese, Mark Hasegawa-Johnson, Odette Scharenborg |
| 2020 | Interspeech | Autosegmental Neural Nets: Should Phones and Tones be Synchronous or Asynchronous? | Jialu Li, Mark Hasegawa-Johnson |
| 2020 | Interspeech | Deep F-Measure Maximization for End-to-End Speech Understanding. | Leda Sari, Mark Hasegawa-Johnson |
| 2020 | Interspeech | A DNN-HMM-DNN Hybrid Model for Discovering Word-Like Units from Spoken Captions and Image Regions. | Liming Wang, Mark Hasegawa-Johnson |
| 2020 | Interspeech | That Sounds Familiar: An Analysis of Phonetic Representations Transfer Across Languages. | Piotr Zelasko, Laureano Moro-Velzquez, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
| 2020 | Interspeech | Identify Speakers in Cocktail Parties with End-to-End Attention. | Junzhe Zhu, Mark Hasegawa-Johnson, Leda Sari |
| 2019 | ICASSP | When CTC Training Meets Acoustic Landmarks. | Di He, Xuesong Yang, Boon Pang Lim, Yi Liang, Mark Hasegawa-Johnson, Deming Chen |
| 2019 | ICASSP | Dimensional Analysis of Laughter in Female Conversational Speech. | Mary Pietrowicz, Carla Agurto, Jonah Casebeer, Mark Hasegawa-Johnson, Karrie Karahalios, Guillermo A. Cecchi |
| 2019 | ICASSP | Pre-training of Speaker Embeddings for Low-latency Speaker Change Detection in Broadcast News. | Leda Sari, Samuel Thomas, Mark Hasegawa-Johnson, Michael Picheny |
| 2019 | ICML | AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss. | Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, Mark Hasegawa-Johnson |
| 2018 | AMIA | Using Conversational Agents to Explain Medication Instructions to Older Adults. | Renato Ferreira Leito Azevedo, Daniel G. Morrow, James F. Graumlich, Ann M. Willemsen-Dunlap, Mark Hasegawa-Johnson, Thomas S. Huang, Kuangxiao Gu, Suma Bhat, Tarek Sakakini, Victor Sadauskas, Donald J. Halpin |
| 2018 | ICASSP | Recognizing Zero-Resourced Languages Based on Mismatched Machine Transcriptions. | Wenda Chen, Mark Hasegawa-Johnson, Nancy F. Chen |
| 2018 | ICASSP | Time-Frequency Networks for Audio Super-Resolution. | Teck-Yian Lim, Raymond A. Yeh, Yijia Xu, Minh N. Do, Mark Hasegawa-Johnson |
| 2018 | ICASSP | Bayesian Models for Unit Discovery on a Very Low Resource Language. | Lucas Ondel, Pierre Godard, Laurent Besacier, Elin Larsen, Mark Hasegawa-Johnson, Odette Scharenborg, Emmanuel Dupoux, Luks Burget, Franois Yvon, Sanjeev Khudanpur |
| 2018 | ICASSP | Deep Learning Based Speech Beamforming. | Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, Dinei A. F. Florncio, Mark Hasegawa-Johnson |
| 2018 | ICASSP | Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop. | Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stker, Pierre Godard, Markus Mller, Lucas Ondel, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang, Emmanuel Dupoux |
| 2018 | ICASSP | Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition. | Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg, Samuel Thomas, Bhuvana Ramabhadran, Mark Hasegawa-Johnson |
| 2018 | ICASSP | Image Restoration with Deep Generative Models. | Raymond A. Yeh, Teck-Yian Lim, Chen Chen, Alexander G. Schwing, Mark Hasegawa-Johnson, Minh N. Do |
| 2018 | Interspeech | Topic and Keyword Identification for Low-resourced Speech Using Cross-Language Transfer Learning. | Wenda Chen, Mark Hasegawa-Johnson, Nancy F. Chen |
| 2018 | Interspeech | Improving DNNs Trained with Non-Native Transcriptions Using Knowledge Distillation and Target Interpolation. | Amit Das, Mark Hasegawa-Johnson |
| 2018 | Interspeech | Improved ASR for Under-resourced Languages through Multi-task Learning with Acoustic Landmarks. | Di He, Boon Pang Lim, Xuesong Yang, Mark Hasegawa-Johnson, Deming Chen |
| 2018 | Interspeech | Speaker Adaptive Audio-Visual Fusion for the Open-Vocabulary Section of AVICAR. | Leda Sari, Mark Hasegawa-Johnson, Kumaran S, Georg Stemmer, Krishnakumar N. Nair |
| 2018 | Interspeech | Visualizing Phoneme Category Adaptation in Deep Neural Networks. | Odette Scharenborg, Sebastian Tiesmeyer, Mark Hasegawa-Johnson, Najim Dehak |
| 2018 | Interspeech | Infant Emotional Outbursts Detection in Infant-parent Spoken Interactions. | Yijia Xu, Mark Hasegawa-Johnson, Nancy McElwain |
| 2017 | ACSSC | Mismatched crowdsourcing: Mining latent skills to acquire speech transcriptions. | Mark Hasegawa-Johnson, Preethi Jyothi, Wenda Chen, Van Hai Do |
| 2017 | AMIA | Using Computer Agents to Explain Clinical Test Results. | Renato Ferreira Leito Azevedo, Kuangxiao Gu, Yang Zhang, Victor Sadauskas, Tarek Sakakini, Daniel G. Morrow, Mark Hasegawa-Johnson, Thomas S. Huang, Suma Bhat, Ann Willemsen-Dunlap, Donald J. Halpin, James F. Graumlich, William Schuh |
| 2017 | AMIA | Dr. Babel Fish: A Machine Translator to Simplify Providers' Language. | Tarek Sakakini, Renato Ferreira Leito Azevedo, Victor Sadauskas, Kuangxiao Gu, Yang Zhang, Suma Bhat, Daniel G. Morrow, Mark Hasegawa-Johnson, Thomas S. Huang, Ann M. Willemsen-Dunlap, Donald J. Halpin, James F. Graumlich |
| 2017 | CVPR | Semantic Image Inpainting with Deep Generative Models. | Raymond A. Yeh, Chen Chen, Teck-Yian Lim, Alexander G. Schwing, Mark Hasegawa-Johnson, Minh N. Do |
| 2017 | ICASSP | Low-resource grapheme-to-phoneme conversion using recurrent neural networks. | Preethi Jyothi, Mark Hasegawa-Johnson |
| 2017 | ICASSP | Discovering dimensions of perceived vocal expression in semi-structured, unscripted oral history accounts. | Mary Pietrowicz, Mark Hasegawa-Johnson, Karrie Karahalios |
| 2017 | Interspeech | Mismatched Crowdsourcing from Multiple Annotator Languages for Recognizing Zero-Resourced Languages: A Nullspace Clustering Approach. | Wenda Chen, Mark Hasegawa-Johnson, Nancy F. Chen, Boon Pang Lim |
| 2017 | Interspeech | Deep Auto-Encoder Based Multi-Task Learning Using Probabilistic Transcriptions. | Amit Das, Mark Hasegawa-Johnson, Karel Vesel |
| 2017 | Interspeech | Multi-Task Learning Using Mismatched Transcription for Under-Resourced Speech Recognition. | Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson |
| 2017 | Interspeech | Using Approximated Auditory Roughness as a Pre-Filtering Feature for Human Screaming and Affective Speech AED. | Di He, Zuofu Cheng, Mark Hasegawa-Johnson, Deming Chen |
| 2017 | Interspeech | Team ELISA System for DARPA LORELEI Speech Evaluation 2016. | Pavlos Papadopoulos, Ruchir Travadi, Colin Vaz, Nikolaos Malandrakis, Ulf Hermjakob, Nima Pourdamghani, Michael Pust, Boliang Zhang, Xiaoman Pan, Di Lu, Ying Lin, Ondrej Glembek, Murali Karthick Baskar, Martin Karafit, Luks Burget, Mark Hasegawa-Johnson, Heng Ji, Jonathan May, Kevin Knight, Shrikanth S. Narayanan |
| 2017 | Interspeech | Speech Enhancement Using Bayesian Wavenet. | Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, Dinei Florncio, Mark Hasegawa-Johnson |
| 2017 | Interspeech | Glottal Model Based Speech Beamforming for ad-hoc Microphone Arrays. | Yang Zhang, Dinei Florncio, Mark Hasegawa-Johnson |
| 2016 | ICASSP | Adapting ASR for under-resourced languages using mismatched transcriptions. | Chunxi Liu, Preethi Jyothi, Hao Tang, Vimal Manohar, Rose Sloan, Tyler Kekona, Mark Hasegawa-Johnson, Sanjeev Khudanpur |
| 2016 | ICASSP | Landmark of Mandarin nasal codas and its application in pronunciation error detection. | Yanlu Xie, Mark Hasegawa-Johnson, Leyuan Qu, Jinsong Zhang |
| 2016 | ICASSP | Stable and symmetric filter convolutional neural network. | Raymond A. Yeh, Mark Hasegawa-Johnson, Minh N. Do |
| 2016 | Interspeech | An Investigation on Training Deep Neural Networks Using Probabilistic Transcriptions. | Amit Das, Mark Hasegawa-Johnson |
| 2016 | Interspeech | Automatic Speech Recognition Using Probabilistic Transcriptions in Swahili, Amharic, and Dinka. | Amit Das, Preethi Jyothi, Mark Hasegawa-Johnson |
| 2016 | Interspeech | Analysis of Mismatched Transcriptions Generated by Humans and Machines for Under-Resourced Languages. | Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson |
| 2016 | ITA | Language coverage for mismatched crowdsourcing. | Lav R. Varshney, Preethi Jyothi, Mark Hasegawa-Johnson |
| 2015 | AAAI | Acquiring Speech Transcriptions Using Mismatched Crowdsourcing. | Preethi Jyothi, Mark Hasegawa-Johnson |
| 2015 | ICASSP | Multichannel transient acoustic signal classification using task-driven dictionary with joint sparsity and beamforming. | Yang Zhang, Nasser M. Nasrabadi, Mark Hasegawa-Johnson |
| 2015 | Interspeech | Cross-lingual transfer learning during supervised training in low resource scenarios. | Amit Das, Mark Hasegawa-Johnson |
| 2015 | Interspeech | Transcribing continuous speech using mismatched crowdsourcing. | Preethi Jyothi, Mark Hasegawa-Johnson |
| 2015 | Interspeech | Improved hindi broadcast ASR by adapting the language model and pronunciation model using a priori syntactic and morphophonemic knowledge. | Preethi Jyothi, Mark Hasegawa-Johnson |
| 2015 | Interspeech | Acoustic correlates for perceived effort levels in expressive speech. | Mary Pietrowicz, Mark Hasegawa-Johnson, Karrie Karahalios |
| 2014 | COLING | A PAC-Bayesian Approach to Minimum Perplexity Language Modeling. | Sujeeth Bharadwaj, Mark Hasegawa-Johnson |
| 2014 | CVPR | Active Planning, Sensing, and Recognition Using a Resource-Constrained Discriminant POMDP. | Zhaowen Wang, Zhangyang Wang, Mark Moll, Po-Sen Huang, Devin K. Grady, Nasser M. Nasrabadi, Thomas S. Huang, Lydia E. Kavraki, Mark Hasegawa-Johnson |
| 2014 | ICASSP | Deep learning for monaural speech separation. | Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, Paris Smaragdis |
| 2014 | ICASSP | Improvement of Probabilistic Acoustic Tube model for speech decomposition. | Yang Zhang, Zhijian Ou, Mark Hasegawa-Johnson |
| 2014 | ICIP | Foreground object detection in highly dynamic scenes using saliency. | Kai-Hsiang Lin, Pooya Khorrami, Jiangping Wang, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2014 | Interspeech | An iterative approach to decision tree training for context dependent speech synthesis. | Xiayu Chen, Yang Zhang, Mark Hasegawa-Johnson |
| 2014 | Interspeech | Detecting articulatory compensation in acoustic data through linear regression modeling. | Alina Khasanova, Jennifer Cole, Mark Hasegawa-Johnson |
| 2014 | LREC | Development of a TV Broadcasts Speech Recognition System for Qatari Arabic. | Mohamed Elmahdy, Mark Hasegawa-Johnson, Eiman Mustafawi |
| 2014 | LREC | Automatic Long Audio Alignment and Confidence Scoring for Conversational Arabic Speech. | Mohamed Elmahdy, Mark Hasegawa-Johnson, Eiman Mustafawi |
| 2013 | ICASSP | Sparse hidden Markov models for purer clusters. | Sujeeth Bharadwaj, Mark Hasegawa-Johnson, Jitendra Ajmera, Om Deshmukh, Ashish Verma |
| 2013 | ICASSP | Random features for Kernel Deep Convex Network. | Po-Sen Huang, Li Deng, Mark Hasegawa-Johnson, Xiaodong He |
| 2013 | ICASSP | Accurate speech segmentation by mimicking human auditory processing. | Sarah King, Mark Hasegawa-Johnson |
| 2012 | COLING | Detection of Acoustic-Phonetic Landmarks in Mismatched Conditions using a Biomimetic Model of Human Auditory Processing. | Sarah King, Mark Hasegawa-Johnson |
| 2012 | ICASSP | Singing-voice separation from monaural recordings using robust principal component analysis. | Po-Sen Huang, Scott Deeann Chen, Paris Smaragdis, Mark Hasegawa-Johnson |
| 2012 | ICASSP | How to put it into words - using random forests to extract symbol level descriptions from audio content for concept detection. | Po-Sen Huang, Robert Mertens, Ajay Divakaran, Gerald Friedland, Mark Hasegawa-Johnson |
| 2012 | ICASSP | Improving faster-than-real-time human acoustic event detection by saliency-maximized audio visualization. | Kai-Hsiang Lin, Xiaodan Zhuang, Camille Goudeseune, Sarah King, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2012 | Interspeech | Pooling Robust Shift-Invariant Sparse Representations of Acoustic Signals. | Po-Sen Huang, Jianchao Yang, Mark Hasegawa-Johnson, Feng Liang, Thomas S. Huang |
| 2012 | Interspeech | F0 and the Perception of Prominence. | Tim Mahrt, Jennifer Cole, Margaret M. Fleck, Mark Hasegawa-Johnson |
| 2011 | FUSION | Multi-sensory features for personnel detection at border crossings. | Po-Sen Huang, Thyagaraju Damarla, Mark Hasegawa-Johnson |
| 2011 | ICASSP | Improving acoustic event detection using generalizable visual features and multi-modality modeling. | Po-Sen Huang, Xiaodan Zhuang, Mark Hasegawa-Johnson |
| 2011 | Interspeech | Optimal Models of Prosodic Prominence Using the Bayesian Information Criterion. | Tim Mahrt, Jui-Ting Huang, Yoonsook Mo, Margaret M. Fleck, Mark Hasegawa-Johnson, Jennifer Cole |
| 2010 | ICASSP | Joint estimation of DOA and speech based on EM beamforming. | Lae-Hoon Kim, Mark Hasegawa-Johnson, Gerasimos Potamianos, Vit Libal |
| 2010 | ICASSP | Toward robust learning of the Gaussian mixture state emission densities for hidden Markov models. | Hao Tang, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2010 | Interspeech | Semi-supervised training of Gaussian mixture models by conditional entropy minimization. | Jui-Ting Huang, Mark Hasegawa-Johnson |
| 2010 | Interspeech | FSM-based pronunciation modeling using articulatory phonological code. | Chi Hu, Xiaodan Zhuang, Mark Hasegawa-Johnson |
| 2010 | Interspeech | Robust automatic speech recognition with decoder oriented ideal binary mask estimation. | Lae-Hoon Kim, Kyung-Tae Kim, Mark Hasegawa-Johnson |
| 2010 | Interspeech | Kinematic analysis of tongue movement control in spastic dysarthria. | Heejin Kim, Panying Rong, Torrey M. Loucks, Mark Hasegawa-Johnson |
| 2010 | Interspeech | A procedure for estimating gestural scores from natural speech. | Hosung Nam, Vikramjit Mitra, Mark Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson, Mark Hasegawa-Johnson |
| 2010 | Interspeech | Landmark-based automated pronunciation error detection. | Su-Youn Yoon, Mark Hasegawa-Johnson, Richard Sproat |
| 2010 | Interspeech | A minimum converted trajectory error (MCTE) approach to high quality speech-to-lips conversion. | Xiaodan Zhuang, Lijuan Wang, Frank K. Soong, Mark Hasegawa-Johnson |
| 2009 | ASRU | Kernel metric learning for phonetic classification. | Jui-Ting Huang, Xi Zhou, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2009 | ICASSP | Acoustic fall detection using Gaussian mixture models and GMM supervectors. | Xiaodan Zhuang, Jing Huang, Gerasimos Potamianos, Mark Hasegawa-Johnson |
| 2009 | Interspeech | Prosodic effects on vowel production: evidence from formant structure. | Yoonsook Mo, Jennifer Cole, Mark Hasegawa-Johnson |
| 2009 | Interspeech | Formant trajectories for acoustic-to-articulatory inversion. | I. Ycel zbek, Mark Hasegawa-Johnson, Mbeccel Demirekler |
| 2009 | Interspeech | Universal access: speech recognition for talkers with spastic dysarthria. | Harsh Vardhan Sharma, Mark Hasegawa-Johnson |
| 2009 | Interspeech | Automated pronunciation scoring using confidence scoring and landmark-based SVM. | Su-Youn Yoon, Mark Hasegawa-Johnson, Richard Sproat |
| 2009 | Interspeech | Articulatory phonological code for word classification. | Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman |
| 2008 | CVPR | Regression from patch-kernel. | Shuicheng Yan, Xi Zhou, Ming Liu, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2008 | ICASSP | Optimal speech estimator considering room response as well as additive noise: Different approaches in low and high frequency range. | Lae-Hoon Kim, Mark Hasegawa-Johnson |
| 2008 | ICASSP | Feature analysis and selection for acoustic event detection. | Xiaodan Zhuang, Xi Zhou, Thomas S. Huang, Mark Hasegawa-Johnson |
| 2008 | ICPR | A novel Gaussianized vector representation for natural scene categorization. | Xi Zhou, Xiaodan Zhuang, Hao Tang, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2008 | ICPR | Face age estimation using patch-based hidden Markov model supervectors. | Xiaodan Zhuang, Xi Zhou, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2008 | Interspeech | Maximum mutual information estimation with unlabeled data for phonetic classification. | Jui-Ting Huang, Mark Hasegawa-Johnson |
| 2008 | Interspeech | Dysarthric speech database for universal access research. | Heejin Kim, Mark Hasegawa-Johnson, Adrienne Perlman, Jon R. Gunderson, Thomas S. Huang, Kenneth L. Watkin, Simone Frame |
| 2008 | Interspeech | Human speech perception and feature extraction. | Bryce E. Lobdell, Mark Hasegawa-Johnson, Jont B. Allen |
| 2008 | Interspeech | Two-stage prosody prediction for emotional text-to-speech synthesis. | Hao Tang, Xi Zhou, Matthias Odisio, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2008 | Interspeech | The entropy of the articulatory phonological code: recognizing gestures from tract variables. | Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis M. Goldstein, Elliot Saltzman |
| 2008 | WACV | EAVA: A 3D Emotive Audio-Visual Avatar. | Hao Tang, Yun Fu, Jilin Tu, Thomas S. Huang, Mark Hasegawa-Johnson |
| 2007 | ICASSP | Articulatory Feature-Based Methods for Acoustic and Audio-Visual Speech Recognition: Summary from the 2006 JHU Summer workshop. | Karen Livescu, zgr etin, Mark Hasegawa-Johnson, Simon King, Chris D. Bartels, Nash M. Borges, Arthur Kantor, Partha Lal, Lisa Yung, Ari Bezman, Stephen Dawson-Haggerty, Bronwyn Woods, Joe Frankel, Mathew Magimai-Doss, Kate Saenko |
| 2007 | ICIP | Lipreading by Locality Discriminant Graph. | Yun Fu, Xi Zhou, Ming Liu, Mark Hasegawa-Johnson, Thomas S. Huang |
| 2007 | Interspeech | Frequency domain correspondence for speaker normalization. | Ming Liu, Xi Zhou, Mark Hasegawa-Johnson, Thomas S. Huang, Zhengyou Zhang |
| 2007 | MMSP | A Multi-Stream Approach to Audiovisual Automatic Speech Recognition. | Mark Hasegawa-Johnson |
| 2006 | ICASSP | Hmm-Based and Svm-Based Recognition of the Speech of Talkers With Spastic Dysarthria. | Mark Hasegawa-Johnson, Jon R. Gunderson, Adrienne Perlman, Thomas S. Huang |
| 2006 | ICASSP | Generalized Optimal Multi-Microphone Speech Enhancement Using Sequential Minimum Variance Distortionless Response(MVDR) Beamforming and Postfiltering. | Lae-Hoon Kim, Mark Hasegawa-Johnson, Koeng-Mo Sung |
| 2006 | Interspeech | Novel time domain multi-class SVMs for landmark detection. | Rahul Chitturi, Mark Hasegawa-Johnson |
| 2006 | Interspeech | Novel entropy based moving average refiners for HMM landmarks. | Rahul Chitturi, Mark Hasegawa-Johnson |
| 2005 | ICASSP | Landmark-Based Speech Recognition: Report of the 2004 Johns Hopkins Summer Workshop. | Mark Hasegawa-Johnson, James Baker, Sarah Borys, Ken Chen, Emily Coogan, Steven Greenberg, Amit Juneja, Katrin Kirchhoff, Karen Livescu, Srividya Mohan, Jennifer Muller, M. Kemal Snmez, Tianyu Wang |
| 2005 | Interspeech | Distinctive feature based SVM discriminant features for improvements to phone recognition on telephone band speech. | Sarah Borys, Mark Hasegawa-Johnson |
| 2004 | ICASSP | An automatic prosody labeling system using ANN-based syntactic-prosodic model and GMM-based acoustic-prosodic model. | Ken Chen, Mark Hasegawa-Johnson, Aaron Cohen |
| 2004 | ICASSP | A factorial HMM approach to simultaneous recognition of isolated digits spoken by multiple talkers on one audio channel. | Ameya N. Deoras, Mark Hasegawa-Johnson |
| 2004 | ICASSP | Formant tracking by mixture state particle filter. | Yanli Zheng, Mark Hasegawa-Johnson |
| 2004 | Interspeech | Modeling and recognition of phonetic and prosodic factors for improvements to acoustic speech recognition models. | Sarah Borys, Aaron Cohen, Mark Hasegawa-Johnson, Jennifer Cole |
| 2004 | Interspeech | Modeling pronunciation variation using artificial neural networks for English spontaneous speech. | Ken Chen, Mark Hasegawa-Johnson |
| 2004 | Interspeech | Source separation using particle filters. | Mital Gandhi, Mark Hasegawa-Johnson |
| 2004 | Interspeech | A factorial HMM aproach to robust isolated digit recognition in background music. | Mark Hasegawa-Johnson, Ameya N. Deoras |
| 2004 | Interspeech | Automatic detection of contrast for speech understanding. | Mark Hasegawa-Johnson, Stephen E. Levinson, Tong Zhang |
| 2004 | Interspeech | Children's emotion recognition in an intelligent tutoring scenario. | Mark Hasegawa-Johnson, Stephen E. Levinson, Tong Zhang |
| 2004 | Interspeech | AVICAR: audio-visual speech corpus in a car environment. | Bowon Lee, Mark Hasegawa-Johnson, Camille Goudeseune, Suketu Kamdar, Sarah Borys, Ming Liu, Thomas S. Huang |
| 2004 | Interspeech | Intertranscriber reliability of prosodic labeling on telephone conversation using toBI. | Taejin Yoon, Sandra Chavarria, Jennifer Cole, Mark Hasegawa-Johnson |
| 2004 | Interspeech | Stop consonant classification by dynamic formant trajectory. | Yanli Zheng, Mark Hasegawa-Johnson, Sarah Borys |
| 2004 | IUI | Semantic analysis for a speech user interface in an intelligent tutoring system. | Yuexi Ren, Mark Hasegawa-Johnson, Stephen E. Levinson |
| 2003 | ICASSP | Acoustic segmentation using switching state Kalman filter. | Yanli Zheng, Mark Hasegawa-Johnson |
| 2003 | Interspeech | Prosody dependent speech recognition with explicit duration modelling at intonational phrase boundaries. | Ken Chen, Sarah Borys, Mark Hasegawa-Johnson, Jennifer Cole |
| 2003 | Interspeech | Maximum conditional mutual information projection for speech recognition. | Mohamed Kamal Omar, Mark Hasegawa-Johnson |
| 2003 | Interspeech | Non-linear maximum likelihood feature transformation for speech recognition. | Mohamed Kamal Omar, Mark Hasegawa-Johnson |
| 2002 | ICASSP | Auditory-modeling inspired methods of feature extraction for robust automatic speech recognition. | Zhinian Jing, Mark Hasegawa-Johnson |
| 2002 | ICASSP | Maximum mutual information based acoustic-features representation of phonological features for speech recognition. | Mohamed Kamal Omar, Mark Hasegawa-Johnson |
| 2002 | Interspeech | An evaluation of using mutual information for selection of acoustic-features representation of phonemes for speech recognition. | Mohamed Kamal Omar, Ken Chen, Mark Hasegawa-Johnson, Yigal Brandman |
| 2001 | ICASSP | PLP coefficients can be quantized at 400 bps. | Wira Gunawan, Mark Hasegawa-Johnson |
| 2000 | ICASSP | Multivariate-state hidden Markov models for simultaneous transcription of phones and formants. | Mark Hasegawa-Johnson |
| 2000 | Interspeech | Time-frequency distribution of partial phonetic information measured using mutual information. | Mark Hasegawa-Johnson |
| 2000 | Interspeech | Signal approximation in Hilbert space and its application on articulatory speech synthesis. | Jun Huang, Stephen E. Levinson, Mark Hasegawa-Johnson |