| 2021 | Unsupervised Cross-Lingual Speech Emotion Recognition Using Pseudo Multilabel. | Jin Li, Nan Yan, Lan Wang |
| 2021 | Improving Speech Recognition on Noisy Speech via Speech Enhancement with Multi-Discriminators CycleGAN. | Chia-Yu Li, Ngoc Thang Vu |
| 2021 | DeepLip: A Benchmark for Deep Learning-Based Audio-Visual Lip Biometrics. | Meng Liu, Longbiao Wang, Kong Aik Lee, Hanyi Zhang, Chang Zeng, Jianwu Dang |
| 2021 | Vibrato Learning in Multi-Singer Singing Voice Synthesis. | Ruolan Liu, Xue Wen, Chunhui Lu, Liming Song, June Sig Sung |
| 2021 | Parameterized Channel Normalization for Far-Field Deep Speaker Verification. | Xuechen Liu, Md. Sahidullah, Tomi Kinnunen |
| 2021 | Optimized Power Normalized Cepstral Coefficients Towards Robust Deep Speaker Verification. | Xuechen Liu, Md. Sahidullah, Tomi Kinnunen |
| 2021 | Utterance-Level Neural Confidence Measure for End-to-End Children Speech Recognition. | Wei Liu, Tan Lee |
| 2021 | Topic Classification on Spoken Documents Using Deep Acoustic and Linguistic Features. | Tan Liu, Wu Guo |
| 2021 | DiffSVC: A Diffusion Probabilistic Model for Singing Voice Conversion. | Songxiang Liu, Yuewen Cao, Dan Su, Helen Meng |
| 2021 | Scaling End-to-End Models for Large-Scale Multilingual ASR. | Bo Li, Ruoming Pang, Tara N. Sainath, Anmol Gulati, Yu Zhang, James Qin, Parisa Haghani, W. Ronny Huang, Min Ma, Junwen Bai |
| 2021 | Uncertainty-Aware Pseudo-Labeling for Spoken Language Assessment. | Binghuai Lin, Liyuan Wang |
| 2021 | Improving Text-Independent Speaker Verification with Auxiliary Speakers Using Graph. | Jingyu Li, Si Ioi Ng, Tan Lee |
| 2021 | SI-Net: Multi-Scale Context-Aware Convolutional Block for Speaker Verification. | Zhuo Li, Ce Fang, Runqiu Xiao, Wenchao Wang, Yonghong Yan |
| 2021 | Improving HS-DACS Based Streaming Transformer ASR with Deep Reinforcement Learning. | Mohan Li, Rama Doddipatla |
| 2021 | Ensemble of Domain Adversarial Neural Networks for Speech Emotion Recognition. | Shi-wook Lee |
| 2021 | Deciding Whether to Ask Clarifying Questions in Large-Scale Spoken Language Understanding. | Joo-Kyung Kim, Guoyin Wang, Sungjin Lee, Young-Bum Kim |
| 2021 | "How Robust R U?": Evaluating Task-Oriented Dialogue Systems on Spoken Conversations. | Seokhwan Kim, Yang Liu, Di Jin, Alexandros Papangelis, Karthik Gopalakrishnan, Behnam Hedayatnia, Dilek Hakkani-Tr |
| 2021 | A Comparison of Streaming Models and Data Augmentation Methods for Robust Speech Recognition. | Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim |
| 2021 | Semi-Supervised Transfer Learning for Language Expansion of End-to-End Speech Recognition Models to Low-Resource Languages. | Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim |
| 2021 | Tiny-CRNN: Streaming Wakeword Detection in a Low Footprint Setting. | Mohammad Omar Khursheed, Christin Jose, Rajath Kumar, Gengshen Fu, Brian Kulis, Santosh Kumar Cheekatmalla |
| 2021 | Learning to Translate Low-Resourced Swiss German Dialectal Speech into Standard German Text. | Abbas Khosravani, Philip N. Garner, Alexandros Lazaridis |
| 2021 | An Evaluation Benchmark for Automatic Speech Recognition of German-English Code-Switching. | Abbas Khosravani, Philip N. Garner, Alexandros Lararidis |
| 2021 | Dialogue Strategy Adaptation to New Action Sets Using Multi-Dimensional Modelling. | Simon Keizer, Norbert Braunschweiler, Svetlana Stoyanchev, Rama Doddipatla |
| 2021 | Attention-Based Multi-Hypothesis Fusion for Speech Summarization. | Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe |
| 2021 | Hybrid Network with Multi-Level Global-Local Statistics Pooling for Robust Text-Independent Speaker Recognition. | Woo Hyun Kang, Jahangir Alam, Abderrahim Fathan |