| 2025 | ASRU | Predictive ASR and Turn-taking Prediction at Once: Towards More Responsive Spoken Dialog System. | Ryo Fukuda, Takatomo Kano, Naohiro Tawara, Marc Delcroix, Atsunori Ogawa, Yuya Chiba, Atsushi Ando |
| 2025 | ASRU | All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR. | Takafumi Moriya, Masato Mimura, Tomohiro Tanaka, Hiroshi Sato, Ryo Masumura, Atsunori Ogawa |
| 2025 | ICASSP | Speech Emotion Recognition Based on Large-Scale Automatic Speech Recognizer. | Ryo Fukuda, Takatomo Kano, Atsushi Ando, Atsunori Ogawa |
| 2025 | ICASSP | Bridging Speech and Text Foundation Models with ReShape Attention. | Takatomo Kano, Atsunori Ogawa, Marc Delcroix, William Chen, Ryo Fukuda, Kohei Matsuura, Takanori Ashihara, Shinji Watanabe |
| 2025 | Interspeech | Why is children's ASR so difficult? Analyzing children's phonological error patterns using SSL-based phoneme recognizers. | Koharu Horii, Naohiro Tawara, Atsunori Ogawa, Shoko Araki |
| 2025 | Interspeech | Pick and Summarize: Integrating Extractive and Abstractive Speech Summarization. | Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Ryo Fukuda, William Chen, Shinji Watanabe |
| 2024 | ICASSP | Train Long and Test Long: Leveraging Full Document Contexts in Speech Processing. | William Chen, Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe |
| 2024 | ICASSP | NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering Diarization. | Naohiro Tawara, Marc Delcroix, Atsushi Ando, Atsunori Ogawa |
| 2024 | Interspeech | Boosting CTC-based ASR using inter-layer attention-based CTC loss. | Keigo Hojo, Yukoh Wakabayashi, Kengo Ohta, Atsunori Ogawa, Norihide Kitaoka |
| 2024 | Interspeech | Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation. | Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Masato Mimura, Takatomo Kano, Atsunori Ogawa, Marc Delcroix |
| 2024 | Interspeech | Text-only Domain Adaptation for CTC-based Speech Recognition through Substitution of Implicit Linguistic Information in the Search Space. | Tatsunari Takagi, Yukoh Wakabayashi, Atsunori Ogawa, Norihide Kitaoka |
| 2023 | ASRU | Summarize While Translating: Universal Model With Parallel Decoding for Summarization and Translation. | Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Kohei Matsuura, Takanori Ashihara, William Chen, Shinji Watanabe |
| 2023 | ASRU | Espnet-Summ: Introducing a Novel Large Dataset, Toolkit, and a Cross-Corpora Evaluation of Speech Summarization Systems. | Roshan S. Sharma, William Chen, Takatomo Kano, Ruchira Sharma, Siddhant Arora, Shinji Watanabe, Atsunori Ogawa, Marc Delcroix, Rita Singh, Bhiksha Raj |
| 2023 | ICASSP | Speech Summarization of Long Spoken Document: Improving Memory Efficiency of Speech/Text Encoders. | Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Roshan S. Sharma, Kohei Matsuura, Shinji Watanabe |
| 2023 | ICASSP | Leveraging Large Text Corpora For End-To-End Speech Summarization. | Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Atsunori Ogawa, Marc Delcroix, Ryo Masumura |
| 2023 | ICASSP | Iterative Shallow Fusion of Backward Language Model for End-To-End Speech Recognition. | Atsunori Ogawa, Takafumi Moriya, Naoyuki Kamo, Naohiro Tawara, Marc Delcroix |
| 2023 | Interspeech | Impact of Residual Noise and Artifacts in Speech Enhancement Errors on Intelligibility of Human and Machine. | Shoko Araki, Ayako Yamamoto, Tsubasa Ochiai, Kenichi Arai, Atsunori Ogawa, Tomohiro Nakatani, Toshio Irino |
| 2023 | Interspeech | Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization. | Marc Delcroix, Naohiro Tawara, Mireia Dez, Federico Landini, Anna Silnova, Atsunori Ogawa, Tomohiro Nakatani, Luks Burget, Shoko Araki |
| 2023 | Interspeech | What are differences? Comparing DNN and Human by Their Performance and Characteristics in Speaker Age Estimation. | Yuki Kitagishi, Naohiro Tawara, Atsunori Ogawa, Ryo Masumura, Taichi Asami |
| 2023 | Interspeech | Transfer Learning from Pre-trained Language Models Improves End-to-End Speech Summarization. | Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Takatomo Kano, Atsunori Ogawa, Marc Delcroix |
| 2023 | Interspeech | Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data. | Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix, Takanori Ashihara, Kohei Matsuura, Tomohiro Tanaka, Ryo Masumura, Atsunori Ogawa, Taichi Asami |
| 2022 | ICASSP | Integrating Multiple ASR Systems into NLP Backend with Attention Fusion. | Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe |
| 2022 | ICASSP | Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models. | Atsunori Ogawa, Naohiro Tawara, Marc Delcroix, Shoko Araki |
| 2022 | Interspeech | End-to-End Spontaneous Speech Recognition Using Disfluency Labeling. | Koharu Horii, Meiko Fukuda, Kengo Ohta, Ryota Nishimura, Atsunori Ogawa, Norihide Kitaoka |
| 2021 | ASRU | Attention-Based Multi-Hypothesis Fusion for Speech Summarization. | Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe |
| 2021 | ASRU | Robust Speech-Age Estimation Using Local Maximum Mean Discrepancy Under Mismatched Recording Conditions. | Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama, Yusuke Ijima |
| 2021 | ICASSP | BLSTM-Based Confidence Estimation for End-to-End Speech Recognition. | Atsunori Ogawa, Naohiro Tawara, Takatomo Kano, Marc Delcroix |
| 2021 | ICASSP | Age-VOX-Celeb: Multi-Modal Corpus for Facial and Speech Estimation. | Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama |
| 2021 | Interspeech | Comparison of Remote Experiments Using Crowdsourcing and Laboratory Experiments on Speech Intelligibility. | Ayako Yamamoto, Toshio Irino, Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani |
| 2020 | ICASSP | Improving Speaker-Attribute Estimation by Voting Based on Speaker Cluster Information. | Naohiro Tawara, Hosana Kamiyama, Satoshi Kobashikawa, Atsunori Ogawa |
| 2020 | ICASSP | Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances. | Naohiro Tawara, Atsunori Ogawa, Tomoharu Iwata, Marc Delcroix, Tetsuji Ogawa |
| 2020 | Interspeech | Predicting Intelligibility of Enhanced Speech Using Posteriors Derived from DNN-Based ASR System. | Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Toshio Irino |
| 2020 | Interspeech | Language Model Data Augmentation Based on Text Domain Transfer. | Atsunori Ogawa, Naohiro Tawara, Marc Delcroix |
| 2019 | ICASSP | A Unified Framework for Feature-based Domain Adaptation of Neural Network Language Models. | Michael Hentschel, Marc Delcroix, Atsunori Ogawa, Tomoharu Iwata, Tomohiro Nakatani |
| 2019 | ICASSP | Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders. | Shigeki Karita, Shinji Watanabe, Tomoharu Iwata, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
| 2019 | ICASSP | A Unified Framework for Neural Speech Separation and Extraction. | Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Tomohiro Nakatani |
| 2019 | ICASSP | ILP-based Compressive Speech Summarization with Content Word Coverage Maximization and Its Oracle Performance Analysis. | Atsunori Ogawa, Tsutomu Hirao, Tomohiro Nakatani, Masaaki Nagata |
| 2019 | Interspeech | Predicting Speech Intelligibility of Enhanced Speech Using Phone Accuracy of DNN-Based ASR System. | Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Katsuhiko Yamamoto, Toshio Irino |
| 2019 | Interspeech | End-to-End SpeakerBeam for Single Channel Target Speech Recognition. | Marc Delcroix, Shinji Watanabe, Tsubasa Ochiai, Keisuke Kinoshita, Shigeki Karita, Atsunori Ogawa, Tomohiro Nakatani |
| 2019 | Interspeech | Improving Transformer-Based End-to-End Speech Recognition with Connectionist Temporal Classification and Language Model Integration. | Shigeki Karita, Nelson Enrique Yalta Soplin, Shinji Watanabe, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
| 2019 | Interspeech | Multimodal SpeakerBeam: Single Channel Target Speech Extraction with Audio-Visual Speaker Clues. | Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Tomohiro Nakatani |
| 2019 | Interspeech | Improved Deep Duel Model for Rescoring N-Best Speech Recognition List Using Backward LSTMLM and Ensemble Encoders. | Atsunori Ogawa, Marc Delcroix, Shigeki Karita, Tomohiro Nakatani |
| 2018 | ICASSP | Single Channel Target Speaker Extraction and Recognition with Speaker Beam. | Marc Delcroix, Katerina Zmolkov, Keisuke Kinoshita, Atsunori Ogawa, Tomohiro Nakatani |
| 2018 | ICASSP | Sequence Training of Encoder-Decoder Model Using Policy Gradient for End-to-End Speech Recognition. | Shigeki Karita, Atsunori Ogawa, Marc Delcroix, Tomohiro Nakatani |
| 2018 | ICASSP | Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations. | Tsuyoshi Morioka, Naohiro Tawara, Tetsuji Ogawa, Atsunori Ogawa, Tomoharu Iwata, Tetsunori Kobayashi |
| 2018 | ICASSP | Rescoring N-Best Speech Recognition List Based on One-on-One Hypothesis Comparison Using Encoder-Classifier Model. | Atsunori Ogawa, Marc Delcroix, Shigeki Karita, Tomohiro Nakatani |
| 2018 | Interspeech | Auxiliary Feature Based Adaptation of End-to-end ASR Systems. | Marc Delcroix, Shinji Watanabe, Atsunori Ogawa, Shigeki Karita, Tomohiro Nakatani |
| 2018 | Interspeech | Semi-Supervised End-to-End Speech Recognition. | Shigeki Karita, Shinji Watanabe, Tomoharu Iwata, Atsunori Ogawa, Marc Delcroix |
| 2017 | ASRU | Learning speaker representation for neural network based multichannel speaker extraction. | Katerina Zmolkov, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani |
| 2017 | ICASSP | Online environmental adaptation of CNN-based acoustic models using spatial diffuseness features. | Christian Huemmer, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Walter Kellermann |
| 2017 | ICASSP | Deep mixture density network for statistical model-based feature enhancement. | Keisuke Kinoshita, Marc Delcroix, Atsunori Ogawa, Takuya Higuchi, Tomohiro Nakatani |
| 2017 | ICASSP | Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models. | Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Taichi Asami, Shigeru Katagiri, Tomohiro Nakatani |
| 2017 | ICASSP | Feedback connection for deep neural network-based acoustic modeling. | Dung T. Tran, Marc Delcroix, Atsunori Ogawa, Christian Huemmer, Tomohiro Nakatani |
| 2017 | Interspeech | Forward-Backward Convolutional LSTM for Acoustic Modeling. | Shigeki Karita, Atsunori Ogawa, Marc Delcroix, Tomohiro Nakatani |
| 2017 | Interspeech | Improved Example-Based Speech Enhancement by Using Deep Neural Network Acoustic Model for Noise Robust Example Search. | Atsunori Ogawa, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani |
| 2017 | Interspeech | Unfolded Deep Recurrent Convolutional Neural Network with Jump Ahead Connections for Acoustic Modeling. | Dung T. Tran, Marc Delcroix, Shigeki Karita, Michael Hentschel, Atsunori Ogawa, Tomohiro Nakatani |
| 2017 | Interspeech | Uncertainty Decoding with Adaptive Sampling for Noise Robust DNN-Based Acoustic Modeling. | Dung T. Tran, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
| 2017 | Interspeech | Speaker-Aware Neural Network Based Beamformer for Speaker Extraction in Speech Mixtures. | Katerina Zmolkov, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani |
| 2016 | ICASSP | Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognition. | Shoko Araki, Masahiro Okada, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani |
| 2016 | ICASSP | Context adaptive deep neural networks for fast acoustic model adaptation in noisy conditions. | Marc Delcroix, Keisuke Kinoshita, Chengzhu Yu, Atsunori Ogawa, Takuya Yoshioka, Tomohiro Nakatani |
| 2016 | Interspeech | Context Adaptive Neural Network for Rapid Adaptation of Deep CNN Based Acoustic Models. | Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Takuya Yoshioka, Dung T. Tran, Tomohiro Nakatani |
| 2016 | Interspeech | Robust Example Search Using Bottleneck Features for Example-Based Speech Enhancement. | Atsunori Ogawa, Shogo Seki, Keisuke Kinoshita, Marc Delcroix, Takuya Yoshioka, Tomohiro Nakatani, Kazuya Takeda |
| 2016 | Interspeech | Factorized Linear Input Network for Acoustic Model Adaptation in Noisy Conditions. | Dung T. Tran, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
| 2015 | ASRU | The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devices. | Takuya Yoshioka, Nobutaka Ito, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Masakiyo Fujimoto, Chengzhu Yu, Wojciech J. Fabian, Miquel Espi, Takuya Higuchi, Shoko Araki, Tomohiro Nakatani |
| 2015 | ICASSP | Double-layer neighborhood graph based similarity search for fast query-by-example spoken term detection. | Kazuo Aoyama, Atsunori Ogawa, Takashi Hattori, Takaaki Hori |
| 2015 | ICASSP | ASR error detection and recognition rate estimation using deep bidirectional recurrent neural networks. | Atsunori Ogawa, Takaaki Hori |
| 2015 | Interspeech | Text-informed speech enhancement with deep neural networks. | Keisuke Kinoshita, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
| 2015 | Interspeech | Robust i-vector extraction for neural network adaptation in noisy environment. | Chengzhu Yu, Atsunori Ogawa, Marc Delcroix, Takuya Yoshioka, Tomohiro Nakatani, John H. L. Hansen |
| 2014 | ICASSP | Zero-resource spoken term detection using hierarchical graph-based similarity search. | Kazuo Aoyama, Atsunori Ogawa, Takashi Hattori, Takaaki Hori, Atsushi Nakamura |
| 2014 | ICASSP | Fast segment search for corpus-based speech enhancement based on speech recognition technology. | Atsunori Ogawa, Keisuke Kinoshita, Takaaki Hori, Tomohiro Nakatani, Atsushi Nakamura |
| 2013 | ICASSP | Graph index based query-by-example search on a large speech data set. | Kazuo Aoyama, Atsunori Ogawa, Takashi Hattori, Takaaki Hori, Atsushi Nakamura |
| 2013 | ICASSP | Unsupervised discriminative adaptation using differenced maximum mutual information based linear regression. | Marc Delcroix, Atsunori Ogawa, Seong-Jun Hahm, Tomohiro Nakatani, Atsushi Nakamura |
| 2013 | ICASSP | Feature space variational Bayesian linear regression and its combination with model space VBLR. | Seong-Jun Hahm, Atsunori Ogawa, Marc Delcroix, Masakiyo Fujimoto, Takaaki Hori, Atsushi Nakamura |
| 2013 | ICASSP | Coupling beamforming with spatial and spectral feature based spectral enhancement and its application to meeting recognition. | Tomohiro Nakatani, Mehrez Souden, Shoko Araki, Takuya Yoshioka, Takaaki Hori, Atsunori Ogawa |
| 2013 | ICASSP | Discriminative recognition rate estimation for N-best list and its application to N-best rescoring. | Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura |
| 2013 | Interspeech | Unsupervised discriminative language modeling using error rate estimator. | Takanobu Oba, Atsunori Ogawa, Takaaki Hori, Hirokazu Masataki, Atsushi Nakamura |
| 2012 | ICASSP | Discriminative feature transforms using differenced maximum mutual information. | Marc Delcroix, Atsunori Ogawa, Shinji Watanabe, Tomohiro Nakatani, Atsushi Nakamura |
| 2012 | ICASSP | Error type classification and word accuracy estimation using alignment features from word confusion network. | Atsunori Ogawa, Takaaki Hori, Atsushi Nakamura |
| 2012 | Interspeech | Speaker Adaptation Using Variational Bayesian Linear Regression in Normalized Feature Space. | Seong-Jun Hahm, Atsunori Ogawa, Masakiyo Fujimoto, Takaaki Hori, Atsushi Nakamura |
| 2012 | Interspeech | Automatic Vocabulary Adaptation Based on Semantic Similarity and Speech Recognition Confidence Measure. | Shoko Yamahata, Yoshikazu Yamaguchi, Atsunori Ogawa, Hirokazu Masataki, Osamu Yoshioka, Satoshi Takahashi |
| 2011 | ICASSP | Machine and acoustical condition dependency analyses for fast acoustic likelihood calculation techniques. | Atsunori Ogawa, Satoshi Takahashi, Atsushi Nakamura |
| 2010 | ICASSP | Discriminative confidence and error cause estimation for extended speech recognition function. | Atsunori Ogawa, Atsushi Nakamura |
| 2010 | Interspeech | A novel confidence measure based on marginalization of jointly estimated error cause probabilities. | Atsunori Ogawa, Atsushi Nakamura |
| 2009 | ICASSP | Efficient combination of likelihood recycling and batch calculation based on conditional fast processing and acoustic back-off. | Atsunori Ogawa, Satoshi Takahashi, Atsushi Nakamura |
| 2009 | Interspeech | Rapid unsupervised adaptation using frame independent output probabilities of gender and context independent phoneme models. | Satoshi Kobashikawa, Atsunori Ogawa, Yoshikazu Yamaguchi, Satoshi Takahashi |
| 2009 | Interspeech | Simultaneous estimation of confidence and error cause in speech recognition using discriminative model. | Atsunori Ogawa, Atsushi Nakamura |
| 2008 | ICASSP | Weighted distance measures for efficient reduction of Gaussian mixture components in HMM-based acoustic model. | Atsunori Ogawa, Satoshi Takahashi |
| 2005 | Interspeech | Rapid response and robust speech recognition by preliminary model adaptation for additive and convolutional noise. | Satoshi Kobashikawa, Satoshi Takahashi, Yoshikazu Yamaguchi, Atsunori Ogawa |
| 2003 | ICASSP | Non-native English speech recognition using bilingual English lexicon and acoustic models. | Shoichi Matsunaga, Atsunori Ogawa, Yoshikazu Yamaguchi, Akihiro Imamura |
| 2003 | Interspeech | Speaker adaptation for non-native speakers using bilingual English lexicon and acoustic models. | Shoichi Matsunaga, Atsunori Ogawa, Yoshikazu Yamaguchi, Akihiro Imamura |
| 2000 | Interspeech | Novel two-pass search strategy using time-asynchronous shortest-first second-pass beam search. | Atsunori Ogawa, Yoshiaki Noda, Shoichi Matsunaga |
| 1998 | ICASSP | Balancing acoustic and linguistic probabilities. | Atsunori Ogawa, Kazuya Takeda, Fumitada Itakura |
| 1998 | Interspeech | Estimating entropy of a language from optimal word insertion penalty. | Kazuya Takeda, Atsunori Ogawa, Fumitada Itakura |