| 2017 | Composite embedding systems for ZeroSpeech2017 Track1. | Hayato Shibata, Taku Kato, Takahiro Shinozaki, Shinji Watanabe |
| 2017 | Topic segmentation in ASR transcripts using bidirectional RNNS for change detection. | Imran A. Sheikh, Dominique Fohr, Irina Illina |
| 2017 | Spoofing detection via simultaneous verification of audio-visual synchronicity and transcription. | Lea Schonherr, Steffen Zeiler, Dorothea Kolossa |
| 2017 | The blizzard machine learning challenge 2017. | Kei Sawada, Keiichi Tokuda, Simon King, Alan W. Black |
| 2017 | Unsupervised adaptation of student DNNS learned from teacher RNNS for improved ASR performance. | Lahiru Samarakoon, Brian Mak |
| 2017 | On lattice generation for large vocabulary speech recognition. | David Rybach, Michael Riley, Johan Schalkwyk |
| 2017 | Simplifying very deep convolutional neural network architectures for robust speech recognition. | Joanna Rownicka, Steve Renals, Peter Bell |
| 2017 | Scalable multi-domain dialogue state tracking. | Abhinav Rastogi, Dilek Hakkani-Tr, Larry P. Heck |
| 2017 | Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer. | Kanishka Rao, Hasim Sak, Rohit Prabhavalkar |
| 2017 | Investigating native and non-native English classification and transfer effects using Legendre polynomial coefficient clustering. | Rachel Rakov, Andrew Rosenberg |
| 2017 | Syllable-based acoustic modeling with CTC-SMBR-LSTM. | Zhongdi Qu, Parisa Haghani, Eugene Weinstein, Pedro J. Moreno |
| 2017 | Exploring ASR-free end-to-end modeling to improve spoken language understanding in a cloud-based dialog system. | Yao Qian, Rutuja Ubale, Vikram Ramanarayanan, Patrick L. Lange, David Suendermann-Oeft, Keelan Evanini, Eugene Tsuprun |
| 2017 | Improving native language (L1) identifation with better VAD and TDNN trained separately on native and non-native English corpora. | Yao Qian, Keelan Evanini, Patrick L. Lange, Robert A. Pugh, Rutuja Ubale, Frank K. Soong |
| 2017 | Deep quaternion neural networks for spoken language understanding. | Titouan Parcollet, Mohamed Morchid, Georges Linars |
| 2017 | A context-aware speech recognition and understanding system for air traffic control domain. | Youssef Oualil, Dietrich Klakow, Gyrgy Szaszk, Ajay Srinivasamurthy, Hartmut Helmke, Petr Motlcek |
| 2017 | Subband wavenet with overlapped single-sideband filterbanks. | Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
| 2017 | Minimally supervised written-to-spoken text normalization. | Axel H. Ng, Kyle Gorman, Richard Sproat |
| 2017 | Consistent DNN uncertainty training and decoding for robust ASR. | Karan Nathwani, Emmanuel Vincent, Irina Illina |
| 2017 | Improving separation of overlapped speech for meeting conversations using uncalibrated microphone array. | Keisuke Nakamura, Randy Gomez |
| 2017 | Automatic speech recognition of Arabic multi-genre broadcast media. | Maryam Najafian, Wei-Ning Hsu, Ahmed Ali, James R. Glass |
| 2017 | Cross-domain speech recognition using nonparallel corpora with cycle-consistent adversarial networks. | Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
| 2017 | Keyword spotting for Google assistant using contextual speech recognition. | Assaf Hurwitz Michaely, Xuedong Zhang, Gabor Simko, Carolina Parada, Petar S. Aleksic |
| 2017 | Binaural processing for robust recognition of degraded speech. | Anjali Menon, Chanwoo Kim, Umpei Kurokawa, Richard M. Stern |
| 2017 | Unsupervised adaptation with domain separation networks for robust speech recognition. | Zhong Meng, Zhuo Chen, Vadim Mazalov, Jinyu Li, Yifan Gong |
| 2017 | Computational cost reduction of long short-term memory based on simultaneous compression of input and hidden state. | Takashi Masuko |