| 2019 | Lead2Gold: Towards Exploiting the Full Potential of Noisy Transcriptions for Speech Recognition. | Adrien Dufraux, Emmanuel Vincent, Awni Y. Hannun, Armelle Brun, Matthijs Douze |
| 2019 | Explicit Alignment of Text and Speech Encodings for Attention-Based End-to-End Speech Recognition. | Jennifer Drexler, James R. Glass |
| 2019 | Improving Fundamental Frequency Generation in EMG-to-Speech Conversion Using a Quantization Approach. | Lorenz Diener, Tejas Umesh, Tanja Schultz |
| 2019 | Optimizing Neural Network Embeddings Using a Pair-Wise Loss for Text-Independent Speaker Verification. | Hira Dhamyal, Tianyan Zhou, Bhiksha Raj, Rita Singh |
| 2019 | SLU for Voice Command in Smart Home: Comparison of Pipeline and End-to-End Approaches. | Thierry Desot, Franois Portet, Michel Vacher |
| 2019 | Long Range Acoustic and Deep Features Perspective on ASVspoof 2019. | Rohan Kumar Das, Jichen Yang, Haizhou Li |
| 2019 | In-the-Wild End-to-End Detection of Speech Affecting Diseases. | M. Joana Correia, Isabel Trancoso, Bhiksha Raj |
| 2019 | Efficient Semi-Supervised Learning for Natural Language Understanding by Optimizing Diversity. | Eunah Cho, He Xie, John P. Lalor, Varun Kumar, William M. Campbell |
| 2019 | A Comparison of End-to-End Models for Long-Form Speech Recognition. | Chung-Cheng Chiu, Anjuli Kannan, Rohit Prabhavalkar, Zhifeng Chen, Tara N. Sainath, Yonghui Wu, Wei Han, Yu Zhang, Ruoming Pang, Sergey Kishchenko, Patrick Nguyen, Arun Narayanan, Hank Liao, Shuyuan Zhang |
| 2019 | Markov Recurrent Neural Network Language Model. | Jen-Tzung Chien, Che-Yu Kuo |
| 2019 | Bayesian Adversarial Learning for Speaker Recognition. | Jen-Tzung Chien, Chun Lin Kuo |
| 2019 | Latent Space Representation for Multi-Target Speaker Detection and Identification with a Sparse Dataset Using Triplet Neural Networks. | Kin Wai Cheuk, Balamurali B. T., Gemma Roig, Dorien Herremans |
| 2019 | Transfer Learning for Context-Aware Spoken Language Understanding. | Qian Chen, Zhu Zhuo, Wen Wang, Qiuyun Xu |
| 2019 | Incremental Lattice Determinization for WFST Decoders. | Zhehuai Chen, Mahsa Yarmohammadi, Hainan Xu, Hang Lv, Lei Xie, Daniel Povey, Sanjeev Khudanpur |
| 2019 | Small-Footprint Keyword Spotting with Graph Convolutional Network. | Xi Chen, Shouyi Yin, Dandan Song, Peng Ouyang, Leibo Liu, Shaojun Wei |
| 2019 | Joint Distribution Learning in the Framework of Variational Autoencoders for Far-Field Speech Enhancement. | Mahesh K. Chelimilla, Shashi Kumar, Shakti P. Rath |
| 2019 | MIMO-Speech: End-to-End Multi-Channel Multi-Speaker Speech Recognition. | Xuankai Chang, Wangyou Zhang, Yanmin Qian, Jonathan Le Roux, Shinji Watanabe |
| 2019 | A Unified Endpointer Using Multitask and Multidomain Training. | Shuo-Yiin Chang, Bo Li, Gabor Simko |
| 2019 | Speaker and Language Aware Training for End-to-End ASR. | Shubham Bansal, Karan Malhotra, Sriram Ganapathy |
| 2019 | Scalable Neural Dialogue State Tracking. | Vevake Balaraman, Bernardo Magnini |
| 2019 | A Comparative Study on End-to-End Speech to Text Translation. | Parnia Bahar, Tobias Bieschke, Hermann Ney |
| 2019 | Learning Hierarchical Representations for Expressive Speaking Style in End-to-End Speech Synthesis. | Xiaochun An, Yuxuan Wang, Shan Yang, Zejun Ma, Lei Xie |
| 2019 | The MGB-5 Challenge: Recognition and Dialect Identification of Dialectal Arabic Speech. | Ahmed Ali, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James R. Glass, Steve Renals, Khalid Choukri |
| 2019 | Novel Enhanced Teager Energy Based Cepstral Coefficients for Replay Spoof Detection. | Rajul Acharya, Hemant A. Patil, Harsh Kotta |
| 2017 | Learning speaker representation for neural network based multichannel speaker extraction. | Katerina Zmolkov, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani |