Tomohiro Tanaka
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
66
Venues
12
Active years
2002–2026
Best venue rank
A*
Where they publish
Papers
66 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2026 | AAAI | Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models. | Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
| 2025 | AAAI | Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level Scores. | Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Satoshi Suzuki, Saki Mizuno, Nobukatsu Hojo |
| 2025 | ASRU | Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model. | Mana Ihori, Taiga Yamane, Naotaka Kawata, Naoki Makishima, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi, Ryo Masumura |
| 2025 | ASRU | Phoneme Overlapping-Aware Pre-Training with External Text Resources for Multi-Talker ASR. | Ryo Masumura, Tomohiro Tanaka, Naoki Makishima, Mana Ihori, Shota Orihashi, Naotaka Kawata, Taiga Yamane, Satoshi Suzuki, Takafumi Moriya |
| 2025 | ASRU | All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR. | Takafumi Moriya, Masato Mimura, Tomohiro Tanaka, Hiroshi Sato, Ryo Masumura, Atsunori Ogawa |
| 2025 | ICAART | Fish Catch Prediction by Combining Fishing, Weather and Tidal Data. | Tomohiro Tanaka, Yasuyuki Tahara, Akihiko Ohsuga, Yuichi Sei |
| 2025 | ICAART | Species-Aware Fish Catch Forecasting Using CatBoost and Multi-source Temporal Features. | Tomohiro Tanaka, Yasuyuki Tahara, Akihiko Ohsuga, Yuichi Sei |
| 2025 | Interspeech | Unified Audio-Visual Modeling for Recognizing Which Face Spoke When and What in Multi-Talker Overlapped Speech and Video. | Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi, Ryo Masumura |
| 2025 | Interspeech | SOMSRED-SVC: Sequential Output Modeling with Speaker Vector Constraints for Joint Multi-Talker Overlapped ASR and Speaker Diarization. | Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi, Ryo Masumura |
| 2024 | ACCV | Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation. | Ryota Yoshihashi, Yuya Otsuka, Kenji Doi, Tomohiro Tanaka, Hirokatsu Kataoka |
| 2024 | ICASSP | Talking Face Generation for Impression Conversion Considering Speech Semantics. | Saki Mizuno, Nobukatsu Hojo, Kazutoshi Shinoda, Keita Suzuki, Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Naotaka Kawata, Satoshi Kobashikawa, Ryo Masumura |
| 2024 | Interspeech | SOMSRED: Sequential Output Modeling for Joint Multi-talker Overlapped Speech Recognition and Speaker Diarization. | Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Atsushi Ando, Ryo Masumura |
| 2024 | Interspeech | Unified Multi-Talker ASR with and without Target-speaker Enrollment. | Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Mana Ihori, Naotaka Kawata, Shota Orihashi, Kazutoshi Shinoda, Taiga Yamane, Saki Mizuno, Keita Suzuki, Satoshi Suzuki, Nobukatsu Hojo, Takafumi Moriya, Atsushi Ando |
| 2023 | ICASSP | Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models. | Takanori Ashihara, Takafumi Moriya, Kohei Matsuura, Tomohiro Tanaka |
| 2023 | ICASSP | Leveraging Large Text Corpora For End-To-End Speech Summarization. | Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Atsunori Ogawa, Marc Delcroix, Ryo Masumura |
| 2023 | ICASSP | Improving Scheduled Sampling for Neural Transducer-Based ASR. | Takafumi Moriya, Takanori Ashihara, Hiroshi Sato, Kohei Matsuura, Tomohiro Tanaka, Ryo Masumura |
| 2023 | ICASSP | Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation Learning. | Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Hiroshi Sato, Taiga Yamane, Takanori Ashihara, Kohei Matsuura, Takafumi Moriya |
| 2023 | ICIP | Ladder Siamese Network: A Method and Insights for Multi-Level Self-Supervised Learning. | Ryota Yoshihashi, Shuhei Nishimura, Dai Yonebayashi, Yuya Otsuka, Tomohiro Tanaka, Takashi Miyazaki |
| 2023 | INLG | Retrieval, Masking, and Generation: Feedback Comment Generation using Masked Comment Examples. | Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Ryo Masumura |
| 2023 | Interspeech | SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge? | Takanori Ashihara, Takafumi Moriya, Kohei Matsuura, Tomohiro Tanaka, Yusuke Ijima, Taichi Asami, Marc Delcroix, Yukinori Honma |
| 2023 | Interspeech | Audio-Visual Praise Estimation for Conversational Video based on Synchronization-Guided Multimodal Transformer. | Nobukatsu Hojo, Saki Mizuno, Satoshi Kobashikawa, Ryo Masumura, Mana Ihori, Hiroshi Sato, Tomohiro Tanaka |
| 2023 | Interspeech | Transcribing Speech as Spoken and Written Dual Text Using an Autoregressive Model. | Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Ryo Masumura, Saki Mizuno, Nobukatsu Hojo |
| 2023 | Interspeech | End-to-End Joint Target and Non-Target Speakers ASR. | Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato, Tomohiro Tanaka, Akihiko Takashima, Satoshi Suzuki, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
| 2023 | Interspeech | Transfer Learning from Pre-trained Language Models Improves End-to-End Speech Summarization. | Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Takatomo Kano, Atsunori Ogawa, Marc Delcroix |
| 2023 | Interspeech | Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data. | Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix, Takanori Ashihara, Kohei Matsuura, Tomohiro Tanaka, Ryo Masumura, Atsunori Ogawa, Taichi Asami |
| 2023 | Interspeech | Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss. | Hiroshi Sato, Ryo Masumura, Tsubasa Ochiai, Marc Delcroix, Takafumi Moriya, Takanori Ashihara, Kentaro Shinayama, Saki Mizuno, Mana Ihori, Tomohiro Tanaka, Nobukatsu Hojo |
| 2022 | COLING | Multi-Perspective Document Revision. | Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Ryo Masumura |
| 2022 | ICASSP | Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration. | Takafumi Moriya, Takanori Ashihara, Atsushi Ando, Hiroshi Sato, Tomohiro Tanaka, Kohei Matsuura, Ryo Masumura, Marc Delcroix, Takahiro Shinozaki |
| 2022 | Interspeech | Deep versus Wide: An Analysis of Student Architectures for Task-Agnostic Knowledge Distillation of Self-Supervised Speech Models. | Takanori Ashihara, Takafumi Moriya, Kohei Matsuura, Tomohiro Tanaka |
| 2022 | Interspeech | End-to-End Joint Modeling of Conversation History-Dependent and Independent ASR Systems with Multi-History Training. | Ryo Masumura, Yoshihiro Yamazaki, Saki Mizuno, Naoki Makishima, Mana Ihori, Mihiro Uchida, Hiroshi Sato, Tomohiro Tanaka, Akihiko Takashima, Satoshi Suzuki, Shota Orihashi, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
| 2022 | Interspeech | Strategies to Improve Robustness of Target Speech Extraction to Enrollment Variations. | Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Takafumi Moriya, Naoki Makishima, Mana Ihori, Tomohiro Tanaka, Ryo Masumura |
| 2022 | Interspeech | Domain Adversarial Self-Supervised Speech Representation Learning for Improving Unknown Domain Downstream Tasks. | Tomohiro Tanaka, Ryo Masumura, Hiroshi Sato, Mana Ihori, Kohei Matsuura, Takanori Ashihara, Takafumi Moriya |
| 2021 | ASRU | Hierarchical Knowledge Distillation for Dialogue Sequence Labeling. | Shota Orihashi, Yoshihiro Yamazaki, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Ryo Masumura |
| 2021 | ICASSP | MAPGN: Masked Pointer-Generator Network for Sequence-to-Sequence Pre-Training. | Mana Ihori, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
| 2021 | ICASSP | Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss. | Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
| 2021 | ICASSP | Hierarchical Transformer-Based Large-Context End-To-End ASR with Large-Context Knowledge Distillation. | Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
| 2021 | ICASSP | Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech Recognition. | Takafumi Moriya, Takanori Ashihara, Tomohiro Tanaka, Tsubasa Ochiai, Hiroshi Sato, Atsushi Ando, Yusuke Ijima, Ryo Masumura, Yusuke Shinohara |
| 2021 | ICDAR | Context-Free TextSpotter for Real-Time and Mobile End-to-End Text Detection and Recognition. | Ryota Yoshihashi, Tomohiro Tanaka, Kenji Doi, Takumi Fujino, Naoaki Yamashita |
| 2021 | Interspeech | Zero-Shot Joint Modeling of Multiple Spoken-Text-Style Conversion Tasks Using Switching Tokens. | Mana Ihori, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
| 2021 | Interspeech | Enrollment-Less Training for Personalized Voice Activity Detection. | Naoki Makishima, Mana Ihori, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
| 2021 | Interspeech | Unified Autoregressive Modeling for Joint End-to-End Multi-Talker Overlapped Speech Recognition and Speaker Attribute Estimation. | Ryo Masumura, Daiki Okamura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
| 2021 | Interspeech | Streaming End-to-End Speech Recognition for Hybrid RNN-T/Attention Architecture. | Takafumi Moriya, Tomohiro Tanaka, Takanori Ashihara, Tsubasa Ochiai, Hiroshi Sato, Atsushi Ando, Ryo Masumura, Marc Delcroix, Taichi Asami |
| 2021 | Interspeech | Cross-Modal Transformer-Based Neural Correction Models for Automatic Speech Recognition. | Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Akihiko Takashima, Takafumi Moriya, Takanori Ashihara, Shota Orihashi, Naoki Makishima |
| 2021 | Interspeech | End-to-End Rich Transcription-Style Automatic Speech Recognition with Semi-Supervised Learning. | Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Akihiko Takashima, Shota Orihashi, Naoki Makishima |
| 2020 | ICASSP | Spoken Language Acquisition Based on Reinforcement Learning and Word Unit Segmentation. | Shengzhou Gao, Wenxin Hou, Tomohiro Tanaka, Takahiro Shinozaki |
| 2020 | ICASSP | Distilling Attention Weights for CTC-Based ASR Systems. | Takafumi Moriya, Hiroshi Sato, Tomohiro Tanaka, Takanori Ashihara, Ryo Masumura, Yusuke Shinohara |
| 2020 | ICPR | Unsupervised Sound Source Localization From Audio-Image Pairs Using Input Gradient Map. | Tomohiro Tanaka, Takahiro Shinozaki |
| 2020 | INLG | Memory Attentive Fusion: External Language Model Integration for Transformer-based Sequence-to-Sequence Model. | Mana Ihori, Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi |
| 2020 | Interspeech | Phoneme-to-Grapheme Conversion Based Large-Scale Pre-Training for End-to-End Automatic Speech Recognition. | Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
| 2020 | Interspeech | Self-Distillation for Improving CTC-Transformer-Based ASR Systems. | Takafumi Moriya, Tsubasa Ochiai, Shigeki Karita, Hiroshi Sato, Tomohiro Tanaka, Takanori Ashihara, Ryo Masumura, Yusuke Shinohara, Marc Delcroix |
| 2020 | Interspeech | Unsupervised Domain Adaptation for Dialogue Sequence Labeling Based on Hierarchical Adversarial Training. | Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Ryo Masumura |
| 2020 | Interspeech | Sound-Image Grounding Based Focusing Mechanism for Efficient Automatic Spoken Language Acquisition. | Mingxin Zhang, Tomohiro Tanaka, Wenxin Hou, Shengzhou Gao, Takahiro Shinozaki |
| 2019 | ASRU | Improving Speech-Based End-of-Turn Detection Via Cross-Modal Representation Learning with Punctuated Text Data. | Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Takanobu Oba, Ryuichiro Higashinaka |
| 2019 | ASRU | Generalized Large-Context Language Models Based on Forward-Backward Hierarchical Recurrent Encoder-Decoder Models. | Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Itsumi Saito, Kyosuke Nishida, Takanobu Oba |
| 2019 | ASRU | Efficient Free Keyword Detection Based on CNN and End-to-End Continuous DP-Matching. | Tomohiro Tanaka, Takahiro Shinozaki |
| 2019 | ICASSP | Large Context End-to-end Automatic Speech Recognition via Extension of Hierarchical Recurrent Encoder-decoder Models. | Ryo Masumura, Tomohiro Tanaka, Takafumi Moriya, Yusuke Shinohara, Takanobu Oba, Yushi Aono |
| 2019 | Interspeech | End-to-End Automatic Speech Recognition with a Reconstruction Criterion Using Speech-to-Text and Text-to-Speech Encoder-Decoders. | Ryo Masumura, Hiroshi Sato, Tomohiro Tanaka, Takafumi Moriya, Yusuke Ijima, Takanobu Oba |
| 2019 | Interspeech | Improving Conversation-Context Language Models with Multiple Spoken Language Understanding Models. | Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Hosana Kamiyama, Takanobu Oba, Satoshi Kobashikawa, Yushi Aono |
| 2019 | Interspeech | Joint Maximization Decoder with Neural Converters for Fully Neural Network-Based Japanese Speech Recognition. | Takafumi Moriya, Jian Wang, Tomohiro Tanaka, Ryo Masumura, Yusuke Shinohara, Yoshikazu Yamaguchi, Yushi Aono |
| 2019 | Interspeech | A Joint End-to-End and DNN-HMM Hybrid Automatic Speech Recognition System with Transferring Sharable Knowledge. | Tomohiro Tanaka, Ryo Masumura, Takafumi Moriya, Takanobu Oba, Yushi Aono |
| 2018 | COLING | Multi-task and Multi-lingual Joint Learning of Neural Lexical Utterance Classification based on Partially-shared Modeling. | Ryo Masumura, Tomohiro Tanaka, Ryuichiro Higashinaka, Hirokazu Masataki, Yushi Aono |
| 2018 | Interspeech | Role Play Dialogue Aware Language Models Based on Conditional Hierarchical Recurrent Encoder-Decoder. | Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Hirokazu Masataki, Yushi Aono |
| 2018 | Interspeech | Neural Error Corrective Language Models for Automatic Speech Recognition. | Tomohiro Tanaka, Ryo Masumura, Hirokazu Masataki, Yushi Aono |
| 2018 | SIGdial | Neural Dialogue Context Online End-of-Turn Detection. | Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Ryuichiro Higashinaka, Yushi Aono |
| 2015 | ASRU | Automation of system building for state-of-the-art large vocabulary speech recognition using evolution strategy. | Takafumi Moriya, Tomohiro Tanaka, Takahiro Shinozaki, Shinji Watanabe, Kevin Duh |
| 2002 | ICASSP | Fundamental frequency estimation based on instantaneous frequency amplitude spectrum. | Tomohiro Tanaka, Takao Kobayashi, Dhany Arifianto, Takashi Masuko |