| 2017 | End-to-end text-independent speaker verification with flexibility in utterance duration. | Chunlei Zhang, Kazuhito Koishida |
| 2017 | Extracting bottleneck features and word-like pairs from untranscribed speech for feature representation. | Yougen Yuan, Cheung-Chi Leung, Lei Xie, Hongjie Chen, Bin Ma, Haizhou Li |
| 2017 | Language diarization for semi-supervised bilingual acoustic model training. | Emre Yilmaz, Mitchell McLaren, Henk van den Heuvel, David A. van Leeuwen |
| 2017 | Noise-robust exemplar matching for rescoring query-by-example search. | Emre Yilmaz, Julien van Hout, Horacio Franco |
| 2017 | Statistical parametric speech synthesis using generative adversarial networks under a multi-task learning framework. | Shan Yang, Lei Xie, Xiao Chen, Xiaoyan Lou, Xuan Zhu, Dongyan Huang, Haizhou Li |
| 2017 | Multi-task ensembles with teacher-student training. | Jeremy Heng Meng Wong, Mark J. F. Gales |
| 2017 | Language independent end-to-end architecture for joint language identification and speech recognition. | Shinji Watanabe, Takaaki Hori, John R. Hershey |
| 2017 | Language modeling with neural trans-dimensional random fields. | Bin Wang, Zhijian Ou |
| 2017 | Integrated speaker-adaptive speech synthesis. | Moquan Wan, Gilles Degottex, Mark J. F. Gales |
| 2017 | Error detection of grapheme-to-phoneme conversion in text-to-speech synthesis using speech signal and lexical context. | Kvin Vythelingum, Yannick Estve, Olivier Rosec |
| 2017 | Denotation extraction for interactive learning in dialogue systems. | Miroslav Vodoln, Filip Jurccek |
| 2017 | MGB-3 but system: Low-resource ASR on Egyptian YouTube data. | Karel Vesel, Murali Karthick Baskar, Mireia Dez, Karel Benes |
| 2017 | Hierarchical recurrent neural network for story segmentation using fusion of lexical and acoustic features. | Emiru Tsunoo, Ondrej Klejch, Peter Bell, Steve Renals |
| 2017 | Attention-based Wav2Text with feature transfer learning. | Andros Tjandra, Sakriani Sakti, Satoshi Nakamura |
| 2017 | Listening while speaking: Speech chain by deep learning. | Andros Tjandra, Sakriani Sakti, Satoshi Nakamura |
| 2017 | Grounded language understanding for manipulation instructions using GAN-based classification. | Komei Sugiura, Hisashi Kawai |
| 2017 | Perceptual quality and modeling accuracy of excitation parameters in DLSTM-based speech synthesis systems. | Eunwoo Song, Frank K. Soong, Hong-Goo Kang |
| 2017 | Reducing the computational complexity for whole word models. | Hagen Soltau, Hank Liao, Hasim Sak |
| 2017 | Aalto system for the 2017 Arabic multi-genre broadcast challenge. | Peter Smit, Siva Reddy Gangireddy, Seppo Enarvi, Sami Virpioja, Mikko Kurimo |
| 2017 | Character-based units for unlimited vocabulary continuous speech recognition. | Peter Smit, Siva Reddy Gangireddy, Seppo Enarvi, Sami Virpioja, Mikko Kurimo |
| 2017 | Improving the efficiency of forward-backward algorithm using batched computation in TensorFlow. | Khe Chai Sim, Arun Narayanan, Tom Bagby, Tara N. Sainath, Michiel Bacchiani |
| 2017 | Leveraging native language speech for accent identification using deep Siamese networks. | Aditya Siddhant, Preethi Jyothi, Sriram Ganapathy |
| 2017 | MIT-QCRI Arabic dialect identification system for the 2017 multi-genre broadcast challenge. | Suwon Shon, Ahmed Ali, James R. Glass |
| 2017 | Multi-view (Joint) probability linear discrimination analysis for J-vector based text dependent speaker verification. | Ziqiang Shi, Liu Liu, Mengjiao Wang, Rujie Liu |
| 2017 | Multitask training with unlabeled data for end-to-end sign language fingerspelling recognition. | Bowen Shi, Karen Livescu |