| 2017 | Adversarial training for data-driven speech enhancement without parallel corpus. | Takuya Higuchi, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani |
| 2017 | Streaming small-footprint keyword spotting using sequence-to-sequence models. | Yanzhang He, Rohit Prabhavalkar, Kanishka Rao, Wei Li, Anton Bakhtin, Ian McGraw |
| 2017 | Feature optimized DPGMM clustering for unsupervised subword modeling: A contribution to zerospeech 2017. | Michael Heck, Sakriani Sakti, Satoshi Nakamura |
| 2017 | An investigation of multi-speaker training for wavenet vocoder. | Tomoki Hayashi, Akira Tamamori, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
| 2017 | Sequence training of DNN acoustic models with natural gradient. | Adnan Haider, Philip C. Woodland |
| 2017 | Semi-supervised training strategies for deep neural networks. | Matthew Gibson, Gary Cook, Puming Zhan |
| 2017 | Investigation of transfer learning for ASR using LF-MMI trained neural networks. | Pegah Ghahremani, Vimal Manohar, Hossein Hadian, Daniel Povey, Sanjeev Khudanpur |
| 2017 | Leveraging side information for speaker identification with the Enron conversational telephone speech collection. | Ning Gao, Gregory Sell, Douglas W. Oard, Mark Dredze |
| 2017 | The zero resource speech challenge 2017. | Ewan Dunbar, Xuan-Nga Cao, Juan Benjumea, Julien Karadayi, Mathieu Bernard, Laurent Besacier, Xavier Anguera, Emmanuel Dupoux |
| 2017 | Exploring the use of acoustic embeddings in neural machine translation. | Salil Deena, Raymond W. M. Ng, Pranava Swaroop Madhyastha, Lucia Specia, Thomas Hain |
| 2017 | Sparse representation of phonetic features for voice conversion with and without parallel data. | Berrak Sisman, Haizhou Li, Kay Chen Tan |
| 2017 | Ground truth estimation of spoken english fluency score using decorrelation penalized low-rank matrix factorization. | Hoon Chung, Yun-Kyung Lee, Jeon Gue Park |
| 2017 | Seeing and hearing too: Audio representation for video captioning. | Shun-Po Chuang, Chia-Hung Wan, Pang-Chi Huang, Chi-Yu Yang, Hung-yi Lee |
| 2017 | Adversarial manifold learning for speaker recognition. | Jen-Tzung Chien, Kang-Ting Peng |
| 2017 | Cracking the cocktail party problem by multi-beam deep attractor network. | Zhuo Chen, Jinyu Li, Xiong Xiao, Takuya Yoshioka, Huaming Wang, Zhenghao Wang, Yifan Gong |
| 2017 | Multilingual bottle-neck feature learning from untranscribed speech. | Hongjie Chen, Cheung-Chi Leung, Lei Xie, Bin Ma, Haizhou Li |
| 2017 | Future word contexts in neural network language models. | Xie Chen, X. Liu, Anton Ragni, Y. Wang, Mark J. F. Gales |
| 2017 | Mitigating the impact of speech recognition errors on chatbot using sequence-to-sequence model. | Pin-Jung Chen, I-Hung Hsu, Yi Yao Huang, Hung-yi Lee |
| 2017 | Dynamic time-aware attention to speaker roles and contexts for spoken language understanding. | Po-Chun Chen, Ta-Chung Chi, Shang-Yu Su, Yun-Nung Chen |
| 2017 | UTD-CRSS submission for MGB-3 Arabic dialect identification: Front-end and back-end advancements on broadcast speech. | Ahmet Emin Bulut, Qian Zhang, Chunlei Zhang, Fahimeh Bahmaninezhad, John H. L. Hansen |
| 2017 | Unwritten languages demand attention too! Word discovery with encoder-decoder models. | Marcely Zanon Boito, Alexandre Berard, Aline Villavicencio, Laurent Besacier |
| 2017 | Exploring neural transducers for end-to-end speech recognition. | Eric Battenberg, Jitong Chen, Rewon Child, Adam Coates, Yashesh Gaur, Yi Li, Hairong Liu, Sanjeev Satheesh, Anuroop Sriram, Zhenyao Zhu |
| 2017 | The CMU entry to blizzard machine learning challenge. | Pallavi Baljekar, Sai Krishna Rallabandi, Alan W. Black |
| 2017 | Meeting recognition with asynchronous distributed microphone array. | Shoko Araki, Nobutaka Ono, Keisuke Kinoshita, Marc Delcroix |
| 2017 | Unsupervised HMM posteriograms for language independent acoustic modeling in zero resource conditions. | T. K. Ansari, Rajath Kumar, Sonali Singh, Sriram Ganapathy, V. Susheela Devi |