Mana Ihori
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
36
Venues
7
Active years
2019–2026
Best venue rank
A*
Where they publish
Papers
36 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2026 | AAAI | Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models. | Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
| 2025 | AAAI | Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level Scores. | Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Satoshi Suzuki, Saki Mizuno, Nobukatsu Hojo |
| 2025 | ASRU | Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model. | Mana Ihori, Taiga Yamane, Naotaka Kawata, Naoki Makishima, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi, Ryo Masumura |
| 2025 | ASRU | Phoneme Overlapping-Aware Pre-Training with External Text Resources for Multi-Talker ASR. | Ryo Masumura, Tomohiro Tanaka, Naoki Makishima, Mana Ihori, Shota Orihashi, Naotaka Kawata, Taiga Yamane, Satoshi Suzuki, Takafumi Moriya |
| 2025 | Interspeech | Unified Audio-Visual Modeling for Recognizing Which Face Spoke When and What in Multi-Talker Overlapped Speech and Video. | Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi, Ryo Masumura |
| 2025 | Interspeech | SOMSRED-SVC: Sequential Output Modeling with Speaker Vector Constraints for Joint Multi-Talker Overlapped ASR and Speaker Diarization. | Naoki Makishima, Naotaka Kawata, Taiga Yamane, Mana Ihori, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi, Ryo Masumura |
| 2024 | ICASSP | Talking Face Generation for Impression Conversion Considering Speech Semantics. | Saki Mizuno, Nobukatsu Hojo, Kazutoshi Shinoda, Keita Suzuki, Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Naotaka Kawata, Satoshi Kobashikawa, Ryo Masumura |
| 2024 | Interspeech | SOMSRED: Sequential Output Modeling for Joint Multi-talker Overlapped Speech Recognition and Speaker Diarization. | Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Atsushi Ando, Ryo Masumura |
| 2024 | Interspeech | Unified Multi-Talker ASR with and without Target-speaker Enrollment. | Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Mana Ihori, Naotaka Kawata, Shota Orihashi, Kazutoshi Shinoda, Taiga Yamane, Saki Mizuno, Keita Suzuki, Satoshi Suzuki, Nobukatsu Hojo, Takafumi Moriya, Atsushi Ando |
| 2023 | ICASSP | Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation Learning. | Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Hiroshi Sato, Taiga Yamane, Takanori Ashihara, Kohei Matsuura, Takafumi Moriya |
| 2023 | INLG | Retrieval, Masking, and Generation: Feedback Comment Generation using Masked Comment Examples. | Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Ryo Masumura |
| 2023 | Interspeech | Audio-Visual Praise Estimation for Conversational Video based on Synchronization-Guided Multimodal Transformer. | Nobukatsu Hojo, Saki Mizuno, Satoshi Kobashikawa, Ryo Masumura, Mana Ihori, Hiroshi Sato, Tomohiro Tanaka |
| 2023 | Interspeech | Transcribing Speech as Spoken and Written Dual Text Using an Autoregressive Model. | Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Ryo Masumura, Saki Mizuno, Nobukatsu Hojo |
| 2023 | Interspeech | End-to-End Joint Target and Non-Target Speakers ASR. | Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato, Tomohiro Tanaka, Akihiko Takashima, Satoshi Suzuki, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
| 2023 | Interspeech | Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss. | Hiroshi Sato, Ryo Masumura, Tsubasa Ochiai, Marc Delcroix, Takafumi Moriya, Takanori Ashihara, Kentaro Shinayama, Saki Mizuno, Mana Ihori, Tomohiro Tanaka, Nobukatsu Hojo |
| 2022 | COLING | Multi-Perspective Document Revision. | Mana Ihori, Hiroshi Sato, Tomohiro Tanaka, Ryo Masumura |
| 2022 | Interspeech | End-to-End Joint Modeling of Conversation History-Dependent and Independent ASR Systems with Multi-History Training. | Ryo Masumura, Yoshihiro Yamazaki, Saki Mizuno, Naoki Makishima, Mana Ihori, Mihiro Uchida, Hiroshi Sato, Tomohiro Tanaka, Akihiko Takashima, Satoshi Suzuki, Shota Orihashi, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
| 2022 | Interspeech | Strategies to Improve Robustness of Target Speech Extraction to Enrollment Variations. | Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Takafumi Moriya, Naoki Makishima, Mana Ihori, Tomohiro Tanaka, Ryo Masumura |
| 2022 | Interspeech | Domain Adversarial Self-Supervised Speech Representation Learning for Improving Unknown Domain Downstream Tasks. | Tomohiro Tanaka, Ryo Masumura, Hiroshi Sato, Mana Ihori, Kohei Matsuura, Takanori Ashihara, Takafumi Moriya |
| 2021 | ASRU | Hierarchical Knowledge Distillation for Dialogue Sequence Labeling. | Shota Orihashi, Yoshihiro Yamazaki, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Ryo Masumura |
| 2021 | ICASSP | MAPGN: Masked Pointer-Generator Network for Sequence-to-Sequence Pre-Training. | Mana Ihori, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
| 2021 | ICASSP | Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss. | Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
| 2021 | ICASSP | Hierarchical Transformer-Based Large-Context End-To-End ASR with Large-Context Knowledge Distillation. | Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
| 2021 | Interspeech | Zero-Shot Joint Modeling of Multiple Spoken-Text-Style Conversion Tasks Using Switching Tokens. | Mana Ihori, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
| 2021 | Interspeech | Enrollment-Less Training for Personalized Voice Activity Detection. | Naoki Makishima, Mana Ihori, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Ryo Masumura |
| 2021 | Interspeech | Unified Autoregressive Modeling for Joint End-to-End Multi-Talker Overlapped Speech Recognition and Speaker Attribute Estimation. | Ryo Masumura, Daiki Okamura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
| 2021 | Interspeech | Cross-Modal Transformer-Based Neural Correction Models for Automatic Speech Recognition. | Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Akihiko Takashima, Takafumi Moriya, Takanori Ashihara, Shota Orihashi, Naoki Makishima |
| 2021 | Interspeech | End-to-End Rich Transcription-Style Automatic Speech Recognition with Semi-Supervised Learning. | Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Akihiko Takashima, Shota Orihashi, Naoki Makishima |
| 2020 | ICASSP | Large-Context Pointer-Generator Networks for Spoken-to-Written Style Conversion. | Mana Ihori, Akihiko Takashima, Ryo Masumura |
| 2020 | ICASSP | Sequence-Level Consistency Training for Semi-Supervised End-to-End Automatic Speech Recognition. | Ryo Masumura, Mana Ihori, Akihiko Takashima, Takafumi Moriya, Atsushi Ando, Yusuke Shinohara |
| 2020 | INLG | Memory Attentive Fusion: External Language Model Integration for Transformer-based Sequence-to-Sequence Model. | Mana Ihori, Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi |
| 2020 | Interspeech | Phoneme-to-Grapheme Conversion Based Large-Scale Pre-Training for End-to-End Automatic Speech Recognition. | Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Shota Orihashi |
| 2020 | Interspeech | Unsupervised Domain Adaptation for Dialogue Sequence Labeling Based on Hierarchical Adversarial Training. | Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Ryo Masumura |
| 2020 | LREC | Parallel Corpus for Japanese Spoken-to-Written Style Conversion. | Mana Ihori, Akihiko Takashima, Ryo Masumura |
| 2019 | ASRU | Improving Speech-Based End-of-Turn Detection Via Cross-Modal Representation Learning with Punctuated Text Data. | Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Takanobu Oba, Ryuichiro Higashinaka |
| 2019 | ASRU | Generalized Large-Context Language Models Based on Forward-Backward Hierarchical Recurrent Encoder-Decoder Models. | Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Itsumi Saito, Kyosuke Nishida, Takanobu Oba |