| 2017 | JHU Kaldi system for Arabic MGB-3 ASR challenge using diarization, audio-transcript alignment and transfer learning. | Vimal Manohar, Daniel Povey, Sanjeev Khudanpur |
| 2017 | A hierarchical attention based model for off-topic spontaneous spoken response detection. | Andrey Malinin, Kate M. Knill, Mark J. F. Gales |
| 2017 | Turbo fusion of magnitude and phase information for DNN-based phoneme recognition. | Timo Lohrenz, Tim Fingscheidt |
| 2017 | Neural relevance-aware query modeling for spoken document retrieval. | Tien-Hong Lo, Ying-Wen Chen, Kuan-Yu Chen, Hsin-Min Wang, Berlin Chen |
| 2017 | Acoustic-to-word model without OOV. | Jinyu Li, Guoli Ye, Rui Zhao, Jasha Droppo, Yifan Gong |
| 2017 | Future vector enhanced LSTM language model for LVCSR. | Qi Liu, Yanmin Qian, Kai Yu |
| 2017 | Iterative policy learning in end-to-end trainable task-oriented neural dialog models. | Bing Liu, Ian R. Lane |
| 2017 | Comparison of multiple features and modeling methods for text-dependent speaker verification. | Yi Liu, Liang He, Yao Tian, Zhuzi Chen, Jia Liu, Michael T. Johnson |
| 2017 | The iFLYTEK system for blizzard machine learning challenge 2017-ES1. | Li-Juan Liu, Chuang Ding, Ya-Jun Hu, Zhen-Hua Ling, Yuan Jiang, Ming Zhou, Si Wei |
| 2017 | Personalized word representations carrying personalized semantics learned from social network posts. | Zih-Wei Lin, Tzu-Wei Sung, Hung-yi Lee, Lin-Shan Lee |
| 2017 | Incremental training and constructing the very deep convolutional residual network acoustic models. | Sheng Li, Xugang Lu, Peng Shen, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
| 2017 | Learning modality-invariant representations for speech and images. | Kenneth Leidal, David Harwath, James R. Glass |
| 2017 | Language modeling with highway LSTM. | Gakuto Kurata, Bhuvana Ramabhadran, George Saon, Abhinav Sethy |
| 2017 | Direct modeling of raw audio with DNNS for wake word detection. | Ken'ichi Kumatani, Sankaran Panchapagesan, Minhua Wu, Minjae Kim, Nikko Strom, Gautam Tiwari, Arindam Mandal |
| 2017 | Lattice rescoring strategies for long short term memory language models in speech recognition. | Shankar Kumar, Michael Nirschl, Daniel Niels Holtmann-Rice, Hank Liao, Ananda Theertha Suresh, Felix X. Yu |
| 2017 | ONENET: Joint domain, intent, slot prediction for spoken language understanding. | Young-Bum Kim, Sungjin Lee, Karl Stratos |
| 2017 | Speaker-sensitive dual memory networks for multi-turn slot tagging. | Young-Bum Kim, Sungjin Lee, Ruhi Sarikaya |
| 2017 | Gated convolutional networks based hybrid acoustic models for low resource speech recognition. | Jian Kang, Wei-Qiang Zhang, Jia Liu |
| 2017 | Investigation of lattice-free maximum mutual information-based acoustic models with sequence-level Kullback-Leibler divergence. | Naoyuki Kanda, Yusuke Fujita, Kenji Nagamatsu |
| 2017 | An embedded segmental K-means model for unsupervised segmentation and clustering of speech. | Herman Kamper, Karen Livescu, Sharon Goldwater |
| 2017 | The USTC system for blizzard machine learning challenge 2017-ES2. | Ya-Jun Hu, Li-Juan Liu, Chuang Ding, Zhen-Hua Ling, Li-Rong Dai |
| 2017 | Unsupervised domain adaptation for robust speech recognition via variational autoencoder-based data augmentation. | Wei-Ning Hsu, Yu Zhang, James R. Glass |
| 2017 | Tackling unseen acoustic conditions in query-by-example search using time and frequency convolution for multilingual deep bottleneck features. | Julien van Hout, Vikramjit Mitra, Horacio Franco, Chris Bartels, Dimitra Vergyri |
| 2017 | Multi-level language modeling and decoding for open vocabulary end-to-end speech recognition. | Takaaki Hori, Shinji Watanabe, John R. Hershey |
| 2017 | Early and late integration of audio features for automatic video description. | Chiori Hori, Takaaki Hori, Tim K. Marks, John R. Hershey |