| 2019 | Joint Optimization of Classification and Clustering for Deep Speaker Embedding. | Zhiming Wang, Kaisheng Yao, Shuo Fang, Xiaolong Li |
| 2019 | Virtual Adversarial Training for DS-CNN Based Small-Footprint Keyword Spotting. | Xiong Wang, Sining Sun, Lei Xie |
| 2019 | Espresso: A Fast End-to-End Neural Speech Recognition Toolkit. | Yiming Wang, Sanjeev Khudanpur, Tongfei Chen, Hainan Xu, Shuoyang Ding, Hang Lv, Yiwen Shao, Nanyun Peng, Lei Xie, Shinji Watanabe |
| 2019 | Using Very Deep Convolutional Neural Networks to Automatically Detect Plagiarized Spoken Responses. | Xinhao Wang, Keelan Evanini, Yao Qian, Klaus Zechner |
| 2019 | Speech Separation Using Speaker Inventory. | Peidong Wang, Zhuo Chen, Xiong Xiao, Zhong Meng, Takuya Yoshioka, Tianyan Zhou, Liang Lu, Jinyu Li |
| 2019 | Dialogue Environments are Different from Games: Investigating Variants of Deep Q-Networks for Dialogue Policy. | Yu-An Wang, Yun-Nung Chen |
| 2019 | Native Language Identification from Raw Waveforms Using Deep Convolutional Neural Networks with Attentive Pooling. | Rutuja Ubale, Vikram Ramanarayanan, Yao Qian, Keelan Evanini, Chee Wee Leong, Chong Min Lee |
| 2019 | Transformer ASR with Contextual Block Processing. | Emiru Tsunoo, Yosuke Kashiwagi, Toshiyuki Kumakura, Shinji Watanabe |
| 2019 | Monotonic Recurrent Neural Network Transducer and Decoding Strategies. | Anshuman Tripathi, Han Lu, Hasim Sak, Hagen Soltau |
| 2019 | Investigation of Shallow Wavenet Vocoder with Laplacian Distribution Output. | Patrick Lumban Tobing, Tomoki Hayashi, Tomoki Toda |
| 2019 | Speech-to-Speech Translation Between Untranscribed Unknown Languages. | Andros Tjandra, Sakriani Sakti, Satoshi Nakamura |
| 2019 | Efficient Free Keyword Detection Based on CNN and End-to-End Continuous DP-Matching. | Tomohiro Tanaka, Takahiro Shinozaki |
| 2019 | Knowledge Distillation from Bert in Pre-Training and Fine-Tuning for Polyphone Disambiguation. | Hao Sun, Xu Tan, Jun-Wei Gan, Sheng Zhao, Dongxu Han, Hongzhi Liu, Tao Qin, Tie-Yan Liu |
| 2019 | Dover: A Method for Combining Diarization Outputs. | Andreas Stolcke, Takuya Yoshioka |
| 2019 | On the Study of Generative Adversarial Networks for Cross-Lingual Voice Conversion. | Berrak Sisman, Mingyang Zhang, Minghui Dong, Haizhou Li |
| 2019 | Emoception: An Inception Inspired Efficient Speech Emotion Recognition Network. | Chirag Singh, Abhay Kumar, Ajay Nagar, Suraj Tripathi, Promod Yenigalla |
| 2019 | Personalization of End-to-End Speech Recognition on Mobile Devices for Named Entities. | Khe Chai Sim, Leif Johnson, Giovanni Motta, Lillian Zhou, Franoise Beaufays, Arnaud Benard, Dhruv Guliani, Andreas Kabel, Nikhil Khare, Tamar Lucassen, Petr Zadrazil, Harry Zhang |
| 2019 | GANs for Children: A Generative Data Augmentation Strategy for Children Speech Recognition. | Peiyao Sheng, Zhuolin Yang, Yanmin Qian |
| 2019 | Attention-Based Speech Recognition Using Gaze Information. | Osamu Segawa, Tomoki Hayashi, Kazuya Takeda |
| 2019 | Simplified LSTMS for Speech Recognition. | George Saon, Zoltn Tske, Kartik Audhkhasi, Brian Kingsbury, Michael Picheny, Samuel Thomas |
| 2019 | Unsupervised Adaptation of Acoustic Models for ASR Using Utterance-Level Embeddings from Squeeze and Excitation Networks. | Hardik B. Sailor, Salil Deena, Md Asif Jalal, Rasa Lileikyte, Thomas Hain |
| 2019 | Embeddings for DNN Speaker Adaptive Training. | Joanna Rownicka, Peter Bell, Steve Renals |
| 2019 | Speech Recognition with Augmented Synthesized Speech. | Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia, Pedro J. Moreno, Yonghui Wu, Zelin Wu |
| 2019 | Paraphrase Generation Based on VAE and Pointer-Generator Networks. | Lohith Ravuru, Hyungtak Choi, Siddarth K. M., Hojung Lee, Inchul Hwang |
| 2019 | Multilingual Bottleneck Features for Query by Example Spoken Term Detection. | Dhananjay Ram, Lesly Miculicich, Herv Bourlard |