| 2021 | Studying Squeeze-and-Excitation Used in CNN for Speaker Verification. | Mickael Rouvier, Pierre-Michel Bousquet |
| 2021 | Comparing the Benefit of Synthetic Training Data for Various Automatic Speech Recognition Architectures. | Nick Rossenbach, Mohammad Zeineldeen, Benedikt Hilmes, Ralf Schlter, Hermann Ney |
| 2021 | AC-VC: Non-Parallel Low Latency Phonetic Posteriorgrams Based Voice Conversion. | Damien Ronssin, Milos Cernak |
| 2021 | Multi-User Voicefilter-Lite via Attentive Speaker Embedding. | Rajeev Rikhye, Quan Wang, Qiao Liang, Yanzhang He, Ian McGraw |
| 2021 | TS-RIR: Translated Synthetic Room Impulse Responses for Speech Augmentation. | Anton Ratnarajah, Zhenyu Tang, Dinesh Manocha |
| 2021 | Conferencingspeech Challenge: Towards Far-Field Multi-Channel Speech Enhancement for Video Conferencing. | Wei Rao, Yihui Fu, Yanxin Hu, Xin Xu, Yvkai Jv, Jiangyu Han, Zhongjie Jiang, Lei Xie, Yannan Wang, Shinji Watanabe, Zheng-Hua Tan, Hui Bu, Tao Yu, Shidong Shang |
| 2021 | Warped Ensembles: A Novel Technique for Improving CTC Based End-to-End Speech Recognition. | Kiran Praveen, Hardik B. Sailor, Abhishek Pandey |
| 2021 | Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch. | Jakob Poncelet, Hugo Van hamme |
| 2021 | Hearing Faces: Target Speaker Text-to-Speech Synthesis from a Face. | Bjrn Plster, Cornelius Weber, Leyuan Qu, Stefan Wermter |
| 2021 | Layer-Wise Analysis of a Self-Supervised Speech Representation Model. | Ankita Pasad, Ju-Chieh Chou, Karen Livescu |
| 2021 | Beyond Isolated Utterances: Conversational Emotion Recognition. | Raghavendra Pappagari, Piotr Zelasko, Jess Villalba, Laureano Moro-Velzquez, Najim Dehak |
| 2021 | Joint Prediction of Truecasing and Punctuation for Conversational Speech in Low-Resource Scenarios. | Raghavendra Pappagari, Piotr Zelasko, Agnieszka Mikolajczyk, Piotr Pezik, Najim Dehak |
| 2021 | Hierarchical Knowledge Distillation for Dialogue Sequence Labeling. | Shota Orihashi, Yoshihiro Yamazaki, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Ryo Masumura |
| 2021 | A Conformer-Based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech Separation. | Tom O'Malley, Arun Narayanan, Quan Wang, Alex Park, James Walker, Nathan Howard |
| 2021 | Multi-Stream HiFi-GAN with Data-Driven Waveform Decomposition. | Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
| 2021 | DEEPA: A Deep Neural Analyzer for Speech and Singing Vocoding. | Sergey Nikonorov, Berrak Sisman, Mingyang Zhang, Haizhou Li |
| 2021 | Cross-Attention Conformer for Context Modeling in Speech Enhancement for ASR. | Arun Narayanan, Chung-Cheng Chiu, Tom O'Malley, Quan Wang, Yanzhang He |
| 2021 | In Pursuit of Babel - Multilingual End-to-End Spoken Language Understanding. | Markus Mller, Samridhi Choudhary, Clement Chung, Athanasios Mouchtaris, Siegfried Kunzmann |
| 2021 | Analysis of Conversational Speech with Application to Voice Adaptation. | Bhagyashree Mukherjee, Anusha Prakash, Hema A. Murthy |
| 2021 | An ASR N-Best Transcript Neural Ranking Model for Spoken Content Retrieval. | Yasufumi Moriya, Gareth J. F. Jones |
| 2021 | Kaizen: Continuously Improving Teacher Using Exponential Moving Average for Semi-Supervised Speech Recognition. | Vimal Manohar, Tatiana Likhomanenko, Qiantong Xu, Wei-Ning Hsu, Ronan Collobert, Yatharth Saraf, Geoffrey Zweig, Abdelrahman Mohamed |
| 2021 | PL-EESR: Perceptual Loss Based End-to-End Robust Speaker Representation Extraction. | Yi Ma, Kong Aik Lee, Ville Hautamki, Haizhou Li |
| 2021 | Multimodal Emotion Recognition with High-Level Speech and Text Features. | Mariana Rodrigues Makiuchi, Kuniaki Uto, Koichi Shinoda |
| 2021 | Audio Embeddings Help to Learn Better Dialogue Policies. | Asier Lpez-Zorrilla, M. Ins Torres, Heriberto Cuayhuitl |
| 2021 | Relaxed Attention: A Simple Method to Boost Performance of End-to-End Automatic Speech Recognition. | Timo Lohrenz, Patrick Schwarz, Zhengyang Li, Tim Fingscheidt |