| 2025 | ASRU | Predictive ASR and Turn-taking Prediction at Once: Towards More Responsive Spoken Dialog System. | Ryo Fukuda, Takatomo Kano, Naohiro Tawara, Marc Delcroix, Atsunori Ogawa, Yuya Chiba, Atsushi Ando |
| 2025 | ASRU | Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization? | Shota Horiguchi, Naohiro Tawara, Takanori Ashihara, Atsushi Ando, Marc Delcroix |
| 2025 | ICASSP | SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model. | Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai, Daisuke Niizumi, Naohiro Tawara, Tomohiro Nakatani, Shoko Araki |
| 2025 | ICASSP | Guided Speaker Embedding. | Shota Horiguchi, Takafumi Moriya, Atsushi Ando, Takanori Ashihara, Hiroshi Sato, Naohiro Tawara, Marc Delcroix |
| 2025 | ICASSP | Mamba-based Segmentation Model for Speaker Diarization. | Alexis Plaquet, Naohiro Tawara, Marc Delcroix, Shota Horiguchi, Atsushi Ando, Shoko Araki |
| 2025 | ICASSP | Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation. | Naohiro Tawara, Atsushi Ando, Shota Horiguchi, Marc Delcroix |
| 2025 | Interspeech | Mitigating Non-Target Speaker Bias in Guided Speaker Embedding. | Shota Horiguchi, Takanori Ashihara, Marc Delcroix, Atsushi Ando, Naohiro Tawara |
| 2025 | Interspeech | Pretraining Multi-Speaker Identification for Neural Speaker Diarization. | Shota Horiguchi, Atsushi Ando, Naohiro Tawara, Marc Delcroix |
| 2025 | Interspeech | Why is children's ASR so difficult? Analyzing children's phonological error patterns using SSL-based phoneme recognizers. | Koharu Horii, Naohiro Tawara, Atsunori Ogawa, Shoko Araki |
| 2024 | ICASSP | Discriminative Training of VBx Diarization. | Dominik Klement, Mireia Dez, Federico Landini, Luks Burget, Anna Silnova, Marc Delcroix, Naohiro Tawara |
| 2024 | ICASSP | NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering Diarization. | Naohiro Tawara, Marc Delcroix, Atsushi Ando, Atsunori Ogawa |
| 2023 | ICASSP | Iterative Shallow Fusion of Backward Language Model for End-To-End Speech Recognition. | Atsunori Ogawa, Takafumi Moriya, Naoyuki Kamo, Naohiro Tawara, Marc Delcroix |
| 2023 | Interspeech | Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization. | Marc Delcroix, Naohiro Tawara, Mireia Dez, Federico Landini, Anna Silnova, Atsunori Ogawa, Tomohiro Nakatani, Luks Burget, Shoko Araki |
| 2023 | Interspeech | What are differences? Comparing DNN and Human by Their Performance and Characteristics in Speaker Age Estimation. | Yuki Kitagishi, Naohiro Tawara, Atsunori Ogawa, Ryo Masumura, Taichi Asami |
| 2023 | Interspeech | Influence of Personal Traits on Impressions of One's Own Voice. | Hikaru Yanagida, Yusuke Ijima, Naohiro Tawara |
| 2022 | ICASSP | Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models. | Atsunori Ogawa, Naohiro Tawara, Marc Delcroix, Shoko Araki |
| 2021 | ASRU | Robust Speech-Age Estimation Using Local Maximum Mean Discrepancy Under Mismatched Recording Conditions. | Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama, Yusuke Ijima |
| 2021 | ICASSP | Integrating End-to-End Neural and Clustering-Based Diarization: Getting the Best of Both Worlds. | Keisuke Kinoshita, Marc Delcroix, Naohiro Tawara |
| 2021 | ICASSP | BLSTM-Based Confidence Estimation for End-to-End Speech Recognition. | Atsunori Ogawa, Naohiro Tawara, Takatomo Kano, Marc Delcroix |
| 2021 | ICASSP | Age-VOX-Celeb: Multi-Modal Corpus for Facial and Speech Estimation. | Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama |
| 2021 | Interspeech | Advances in Integration of End-to-End Neural and Clustering-Based Diarization for Real Conversational Speech. | Keisuke Kinoshita, Marc Delcroix, Naohiro Tawara |
| 2020 | ICASSP | Improving Speaker Discrimination of Target Speech Extraction With Time-Domain Speakerbeam. | Marc Delcroix, Tsubasa Ochiai, Katerina Zmolkov, Keisuke Kinoshita, Naohiro Tawara, Tomohiro Nakatani, Shoko Araki |
| 2020 | ICASSP | Improving Speaker-Attribute Estimation by Voting Based on Speaker Cluster Information. | Naohiro Tawara, Hosana Kamiyama, Satoshi Kobashikawa, Atsunori Ogawa |
| 2020 | ICASSP | Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances. | Naohiro Tawara, Atsunori Ogawa, Tomoharu Iwata, Marc Delcroix, Tetsuji Ogawa |
| 2020 | Interspeech | Language Model Data Augmentation Based on Text Domain Transfer. | Atsunori Ogawa, Naohiro Tawara, Marc Delcroix |
| 2019 | ICASSP | Postfiltering Using an Adversarial Denoising Autoencoder with Noise-aware Training. | Naohiro Tawara, Hikari Tanabe, Tetsunori Kobayashi, Masaru Fujieda, Kazuhiro Katagiri, Takashi Yazu, Tetsuji Ogawa |
| 2019 | Interspeech | Speaker Adversarial Training of DPGMM-Based Feature Extractor for Zero-Resource Languages. | Yosuke Higuchi, Naohiro Tawara, Tetsunori Kobayashi, Tetsuji Ogawa |
| 2019 | Interspeech | Multi-Channel Speech Enhancement Using Time-Domain Convolutional Denoising Autoencoder. | Naohiro Tawara, Tetsunori Kobayashi, Tetsuji Ogawa |
| 2018 | ICASSP | Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations. | Tsuyoshi Morioka, Naohiro Tawara, Tetsuji Ogawa, Atsunori Ogawa, Tomoharu Iwata, Tetsunori Kobayashi |
| 2018 | ICASSP | Speaker Invariant Feature Extraction for Zero-Resource Languages with Adversarial Learning. | Taira Tsuchiya, Naohiro Tawara, Tetsuji Ogawa, Tetsunori Kobayashi |
| 2018 | ICPR | Sequential Fish Catch Forecasting Using Bayesian State Space Models. | Yuya Kokaki, Naohiro Tawara, Tetsunori Kobayashi, Kazuo Hashimoto, Tetsuji Ogawa |
| 2015 | ICASSP | A comparative study of spectral clustering for i-vector-based speaker clustering under noisy conditions. | Naohiro Tawara, Tetsuji Ogawa, Tetsunori Kobayashi |
| 2012 | ICASSP | Fully Bayesian inference of multi-mixture Gaussian model and its evaluation using speaker clustering. | Naohiro Tawara, Tetsuji Ogawa, Shinji Watanabe, Tetsunori Kobayashi |
| 2012 | Interspeech | Fully Bayesian speaker clustering based on hierarchically structured utterance-oriented Dirichlet process mixture model. | Naohiro Tawara, Tetsuji Ogawa, Shinji Watanabe, Atsushi Nakamura, Tetsunori Kobayashi |
| 2011 | Interspeech | Speaker Clustering Based on Utterance-Oriented Dirichlet Process Mixture Model. | Naohiro Tawara, Shinji Watanabe, Tetsuji Ogawa, Tetsunori Kobayashi |