| 2025 | ASRU | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities. | George Saon, Avihu Dekel, Alexander Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Xiaodong Cui, Zvi Kons |
| 2025 | ICASSP | Knowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASR. | Takashi Fukuda, Gakuto Kurata, George Saon |
| 2025 | ICASSP | LLM based Text Generation for Improved Low-resource Speech Recognition Models. | Tohru Nagano, Gakuto Kurata, Samuel Thomas, Hong-Kwang Jeff Kuo, Daniel Bolaos, Hyun Jung, George Saon |
| 2025 | Interspeech | Improving End-to-end Mixed-case ASR with Knowledge Distillation and Integration of Voice Activity Cues. | Sashi Novitasari, Takashi Fukuda, Gakuto Kurata |
| 2025 | Interspeech | Voice Activity-based Text Segmentation for ASR Text Denormalization. | Sashi Novitasari, Takashi Fukuda, Gakuto Kurata |
| 2024 | EMNLP | Robust ASR Error Correction with Conservative Data Filtering. | Takuma Udagawa, Masayuki Suzuki, Masayasu Muraoka, Gakuto Kurata |
| 2024 | ICASSP | Multiple Representation Transfer from Large Language Models to End-to-End ASR Systems. | Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Masayasu Muraoka, George Saon |
| 2023 | EMNLP | Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word Dictionaries. | Ashish R. Mittal, Sunita Sarawagi, Preethi Jyothi, George Saon, Gakuto Kurata |
| 2022 | Interspeech | Improving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label Smoothing. | Xiaodong Cui, George Saon, Tohru Nagano, Masayuki Suzuki, Takashi Fukuda, Brian Kingsbury, Gakuto Kurata |
| 2022 | Interspeech | Global RNN Transducer Models For Multi-dialect Speech Recognition. | Takashi Fukuda, Samuel Thomas, Masayuki Suzuki, Gakuto Kurata, George Saon, Brian Kingsbury |
| 2022 | Interspeech | Improving ASR Robustness in Noisy Condition Through VAD Integration. | Sashi Novitasari, Takashi Fukuda, Gakuto Kurata |
| 2022 | Interspeech | Effect and Analysis of Large-scale Language Model Rescoring on Competitive ASR Systems. | Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Nobuyasu Itoh, George Saon |
| 2021 | ICASSP | RNN Transducer Models for Spoken Language Understanding. | Samuel Thomas, Hong-Kwang Jeff Kuo, George Saon, Zoltn Tske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory |
| 2021 | ICASSP | Generalized Knowledge Distillation from an Ensemble of Specialized Teachers Leveraging Unsupervised Neural Clustering. | Takashi Fukuda, Gakuto Kurata |
| 2021 | Interspeech | Improving Customization of Neural Transducers by Mitigating Acoustic Mismatch of Synthesized Audio. | Gakuto Kurata, George Saon, Brian Kingsbury, David Haws, Zoltn Tske |
| 2020 | ICASSP | Converting Written Language to Spoken Language with Neural Machine Translation for Language Modeling. | Shintaro Ando, Masayuki Suzuki, Nobuyasu Itoh, Gakuto Kurata, Nobuaki Minematsu |
| 2020 | ICASSP | Speaker Embeddings Incorporating Acoustic Conditions for Diarization. | Yosuke Higuchi, Masayuki Suzuki, Gakuto Kurata |
| 2020 | Interspeech | New Advances in Speaker Diarization. | Hagai Aronowitz, Weizhong Zhu, Masayuki Suzuki, Gakuto Kurata, Ron Hoory |
| 2020 | Interspeech | End-to-End Spoken Language Understanding Without Full Transcripts. | Hong-Kwang Jeff Kuo, Zoltn Tske, Samuel Thomas, Yinghui Huang, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, Luis A. Lastras |
| 2020 | Interspeech | Knowledge Distillation from Offline to Streaming RNN Transducer for End-to-End Speech Recognition. | Gakuto Kurata, George Saon |
| 2019 | ASRU | Data Augmentation Based on Vowel Stretch for Improving Children's Speech Recognition. | Tohru Nagano, Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata |
| 2019 | ICASSP | Improvements to N-gram Language Model Using Text Generated from Neural Language Model. | Masayuki Suzuki, Nobuyasu Itoh, Tohru Nagano, Gakuto Kurata, Samuel Thomas |
| 2019 | ICASSP | English Broadcast News Speech Recognition by Humans and Machines. | Samuel Thomas, Masayuki Suzuki, Yinghui Huang, Gakuto Kurata, Zoltn Tske, George Saon, Brian Kingsbury, Michael Picheny, Tom Dibert, Alice Kaiser-Schatzlein, Bern Samko |
| 2019 | Interspeech | Direct Neuron-Wise Fusion of Cognate Neural Networks. | Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata |
| 2019 | Interspeech | Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge Distillation. | Gakuto Kurata, Kartik Audhkhasi |
| 2019 | Interspeech | Multi-Task CTC Training with Auxiliary Feature Reconstruction for End-to-End Speech Recognition. | Gakuto Kurata, Kartik Audhkhasi |
| 2018 | Interspeech | Data Augmentation Improves Recognition of Foreign Accented Speech. | Takashi Fukuda, Raul Fernandez, Andrew Rosenberg, Samuel Thomas, Bhuvana Ramabhadran, Alexander Sorin, Gakuto Kurata |
| 2018 | Interspeech | Inference-Invariant Transformation of Batch Normalization for Domain Adaptation of Acoustic Models. | Masayuki Suzuki, Tohru Nagano, Gakuto Kurata, Samuel Thomas |
| 2017 | ASRU | Language modeling with highway LSTM. | Gakuto Kurata, Bhuvana Ramabhadran, George Saon, Abhinav Sethy |
| 2017 | ICASSP | Effective joint training of denoising feature space transforms and Neural Network based acoustic models. | Takashi Fukuda, Osamu Ichikawa, Gakuto Kurata, Ryuki Tachibana, Samuel Thomas, Bhuvana Ramabhadran |
| 2017 | ICASSP | Harmonic feature fusion for robust neural network-based acoustic modeling. | Osamu Ichikawa, Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Bhuvana Ramabhadran |
| 2017 | Interspeech | Efficient Knowledge Distillation from an Ensemble of Teachers. | Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas, Jia Cui, Bhuvana Ramabhadran |
| 2017 | Interspeech | Ensembles of Multi-Scale VGG Acoustic Models. | Michael Heck, Masayuki Suzuki, Takashi Fukuda, Gakuto Kurata, Satoshi Nakamura |
| 2017 | Interspeech | Factorial Modeling for Effective Suppression of Directional Noise. | Osamu Ichikawa, Takashi Fukuda, Gakuto Kurata, Steven J. Rennie |
| 2017 | Interspeech | Empirical Exploration of Novel Architectures and Objectives for Language Models. | Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, George Saon |
| 2017 | Interspeech | English Conversational Telephone Speech Recognition by Humans and Machines. | George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas, Dimitrios Dimitriadis, Xiaodong Cui, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
| 2017 | Interspeech | Symbol Sequence Search from Telephone Conversation. | Masayuki Suzuki, Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, Kenneth Ward Church, Mark Drake |
| 2016 | EMNLP | Leveraging Sentence-level Information with Encoder LSTM for Semantic Slot Filling. | Gakuto Kurata, Bing Xiang, Bowen Zhou, Mo Yu |
| 2016 | ICASSP | Speech recognition robust against speech overlapping in monaural recordings of telephone conversations. | Masayuki Suzuki, Gakuto Kurata, Tohru Nagano, Ryuki Tachibana |
| 2016 | Interspeech | Improved Neural Network Initialization by Grouping Context-Dependent Targets for Acoustic Modeling. | Gakuto Kurata, Brian Kingsbury |
| 2016 | Interspeech | Labeled Data Generation with Encoder-Decoder LSTM for Semantic Slot Filling. | Gakuto Kurata, Bing Xiang, Bowen Zhou |
| 2016 | NAACL | Improved Neural Network-based Multi-label Classification with Better Initialization Leveraging Label Co-occurrence. | Gakuto Kurata, Bing Xiang, Bowen Zhou |
| 2015 | Interspeech | A metric for evaluating speech recognizer output based on human-perception model. | Nobuyasu Itoh, Gakuto Kurata, Ryuki Tachibana, Masafumi Nishimura |
| 2015 | Interspeech | Deep neural network training emphasizing central frames. | Gakuto Kurata, Daniel Willett |
| 2012 | Interspeech | Discriminative Reranking for LVCSR Leveraging Invariant Structure. | Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
| 2011 | ICASSP | Training of error-corrective model for ASR without using audio data. | Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura |
| 2011 | ICASSP | Named entity recognition from Conversational Telephone Speech leveraging Word Confusion Networks for training and recognition. | Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura, Abhinav Sethy, Bhuvana Ramabhadran |
| 2011 | Interspeech | Acoustic Model Training with Detecting Transcription Errors in the Training Data. | Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura |
| 2011 | Interspeech | Continuous Digits Recognition Leveraging Invariant Structure. | Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
| 2009 | ICASSP | Acoustically discriminative training for language models. | Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura |
| 2007 | ICASSP | Unsupervised Lexicon Acquisition from Speech and Text. | Gakuto Kurata, Shinsuke Mori, Nobuyasu Itoh, Masafumi Nishimura |
| 2007 | Interspeech | Preliminary experiments toward automatic generation of new TTS voices from recorded speech alone. | Ryuki Tachibana, Tohru Nagano, Gakuto Kurata, Masafumi Nishimura, Noboru Babaguchi |
| 2006 | ACL | Phoneme-to-Text Transcription System with an Infinite Vocabulary. | Shinsuke Mori, Daisuke Takuma, Gakuto Kurata |
| 2006 | ICASSP | Unsupervised Adaptation of a Stochastic Language Model Using a Japanese Raw Corpus. | Gakuto Kurata, Shinsuke Mori, Masafumi Nishimura |
| 2005 | Interspeech | Class-based variable memory length Markov model. | Shinsuke Mori, Gakuto Kurata |
| 2002 | Interspeech | Integration of MLLR adaptation with pronunciation proficiency adaptation for non-native speech recognition. | Nobuaki Minematsu, Gakuto Kurata, Keikichi Hirose |
| 2002 | Interspeech | Corpus-based analysis of English spoken by Japanese students in view of the entire phonemic system of English. | Nobuaki Minematsu, Gakuto Kurata, Keikichi Hirose |