| 2026 | HRI | Preliminary Evaluation of Multimodal Expressions by a Reception Robot in Three Situations. | Carlos Toshinori Ishi, Taiki Yano, Yuka Nakayama |
| 2025 | ICMI | SignFlow: End-to-End Sign Language Generation for One-to-Many Modeling using Conditional Flow Matching. | Nabeela Khan, Bowen Wu, Sihan Tan, Carlos Toshinori Ishi, Kazuhiro Nakadai |
| 2025 | Interspeech | What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems. | Kiyotada Mori, Seiya Kawano, Chaoran Liu, Carlos Toshinori Ishi, Angel F. Garcia Contreras, Koichiro Yoshino |
| 2025 | MMM | RoboDJ: Live Commentary Robots System Driven by Physical- and Cyber-World Observations. | Yasutomo Kawanishi, Yutaka Nakamura, Taiken Shintani, Carlos Toshinori Ishi, Seiya Kawano, Koichiro Yoshino, Takashi Minato, Michihiko Minoh |
| 2024 | Interspeech | X-E-Speech: Joint Training Framework of Non-Autoregressive Cross-lingual Emotional Text-to-Speech and Voice Conversion. | Houjian Guo, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2024 | IROS | Retargeting Human Facial Expression to Human-like Robotic Face through Neural Network Surrogate-based Optimization. | Bowen Wu, Chaoran Liu, Carlos Toshinori Ishi, Takashi Minato, Hiroshi Ishiguro |
| 2024 | RO-MAN | Age and Spatial Cue Effects on User Performance for an Adaptable Verbal Wayfinding System. | Rin Takahira, Chaoran Liu, Carlos Toshinori Ishi, Takenao Ohkawa |
| 2023 | ASRU | QUICKVC: A Lightweight VITS-Based Any-to-Many Voice Conversion Model using ISTFT for Faster Conversion. | Houjian Guo, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2023 | ASRU | Using Joint Training Speaker Encoder With Consistency Loss to Achieve Cross-Lingual Voice Conversion and Expressive Voice Conversion. | Houjian Guo, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2023 | CHI | I Know Your Feelings Before You Do: Predicting Future Affective Reactions in Human-Computer Dialogue. | Yuanchao Li, Koji Inoue, Leimin Tian, Changzeng Fu, Carlos Toshinori Ishi, Hiroshi Ishiguro, Tatsuya Kawahara, Catherine Lai |
| 2023 | ICASSP | HAG: Hierarchical Attention with Graph Network for Dialogue Act Classification in Conversation. | Changzeng Fu, Zhenghan Chen, Jiaqi Shi, Bowen Wu, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2023 | IROS | Recognizing Real-World Intentions using A Multimodal Deep Learning Approach with Spatial-Temporal Graph Convolutional Networks. | Jiaqi Shi, Chaoran Liu, Carlos Toshinori Ishi, Bowen Wu, Hiroshi Ishiguro |
| 2022 | ACII | A Controllable Cross-Gender Voice Conversion for Social Robot. | Changzeng Fu, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2022 | HRI | Butsukusa: A Conversational Mobile Robot Describing Its Own Observations and Internal States. | Akishige Yuguchi, Seiya Kawano, Koichiro Yoshino, Carlos Toshinori Ishi, Yasutomo Kawanishi, Yutaka Nakamura, Takashi Minato, Yasuki Saito, Michihiko Minoh |
| 2022 | IROS | Controlling the Impression of Robots via GAN-based Gesture Generation. | Bowen Wu, Jiaqi Shi, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2022 | RO-MAN | Expression of Personality by Gaze Movements of an Android Robot in Multi-Party Dialogues | Taiken Shintani, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2021 | HAI | Analysis of Role-Based Gaze Behaviors and Gaze Aversions, and Implementation of Robot's Gaze Control for Multi-party Dialogue. | Taiken Shintani, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2021 | ICASSP | MAEC: Multi-Instance Learning with an Adversarial Auto-Encoder-Based Classifier for Speech Emotion Recognition. | Changzeng Fu, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2021 | ICMI | Probabilistic Human-like Gesture Synthesis from Speech using GRU-based WGAN. | Bowen Wu, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2021 | Interspeech | Analysis of Eye Gaze Reasons and Gaze Aversions During Three-Party Conversations. | Carlos Toshinori Ishi, Taiken Shintani |
| 2020 | HRI | Generation and Evaluation of Audio-Visual Anger Emotional Expression for Android Robot. | Chinenye Augustine Ajibo, Ryusuke Mikata, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2019 | Interspeech | A Neural Turn-Taking Model without RNN. | Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2019 | RO-MAN | Analysis of factors influencing the impression of speaker individuality in android robots. | Ryusuke Mikata, Carlos Toshinori Ishi, Takashi Minato, Hiroshi Ishiguro |
| 2017 | Interspeech | Prosodic Analysis of Attention-Drawing Speech. | Carlos Toshinori Ishi, Jun Arai, Norihiro Hagita |
| 2017 | Interspeech | Motion Analysis in Vocalized Surprise Expressions. | Carlos Toshinori Ishi, Takashi Minato, Hiroshi Ishiguro |
| 2017 | Interspeech | Turn-Taking Estimation Model Based on Joint Embedding of Lexical and Prosodic Contents. | Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2017 | IROS | Probabilistic nod generation model based on estimated utterance categories. | Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2016 | IROS | Motion generation in android robots during laughing speech. | Carlos Toshinori Ishi, Tomo Funayama, Takashi Minato, Hiroshi Ishiguro |
| 2016 | IROS | Hearing support system using environment sensor network. | Carlos Toshinori Ishi, Chaoran Liu, Jani Even, Norihiro Hagita |
| 2016 | RO-MAN | ERICA: The ERATO Intelligent Conversational Android. | Dylan F. Glas, Takashi Minato, Carlos Toshinori Ishi, Tatsuya Kawahara, Hiroshi Ishiguro |
| 2016 | RO-MAN | Speech driven trunk motion generating system based on physical constraint. | Kurima Sakai, Takashi Minato, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2015 | HRI | Bringing the Scene Back to the Tele-operator: Auditory Scene Manipulation for Tele-presence Systems. | Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2015 | IROS | Audio augmented point clouds for applications in robotics. | Jani Even, Florent Ferreri, Atsushi Watanabe, Yoichi Morales, Carlos Toshinori Ishi, Norihiro Hagita |
| 2015 | IROS | Speech activity detection and face orientation estimation using multiple microphone arrays and human position information. | Carlos Toshinori Ishi, Jani Even, Norihiro Hagita |
| 2015 | IROS | Robot-assisted acoustic inspection of infrastructures - cooperative hammer sounding inspection. | Atsushi Watanabe, Jani Even, Luis Yoichi Morales Saiki, Carlos Toshinori Ishi |
| 2015 | RO-MAN | Online speech-driven head motion generating system and evaluation on a tele-operated robot. | Kurima Sakai, Carlos Toshinori Ishi, Takashi Minato, Hiroshi Ishiguro |
| 2014 | ICRA | Mapping sound emitting structures in 3D. | Jani Even, Yoichi Morales, Nagasrikanth Kallakuri, Jonas Furrer, Carlos Toshinori Ishi, Norihiro Hagita |
| 2014 | Interspeech | Analysis of laughter events in real science classes by using multiple environment sensor data. | Carlos Toshinori Ishi, Hiroaki Hatano, Norihiro Hagita |
| 2014 | IROS | Audio ray tracing for position estimation of entities in blind regions. | Jani Even, Yoichi Morales, Nagasrikanth Kallakuri, Carlos Toshinori Ishi, Norihiro Hagita |
| 2013 | ICRA | Probabilistic approach for building auditory maps with a mobile microphone array. | Nagasrikanth Kallakuri, Jani Even, Yoichi Morales Saiki, Carlos Toshinori Ishi, Norihiro Hagita |
| 2013 | Interspeech | Analysis of factors involved in the choice of rising or non-rising intonation in question utterances appearing in conversational speech. | Hiroaki Hatano, Miyako Kiso, Carlos Toshinori Ishi |
| 2013 | IROS | Creation of radiated sound intensity maps using multi-modal measurements onboard an autonomous mobile platform. | Jani Even, Nagasrikanth Kallakuri, Yoichi Morales Saiki, Carlos Toshinori Ishi, Norihiro Hagita |
| 2013 | IROS | Using multiple microphone arrays and reflections for 3D localization of sound sources. | Carlos Toshinori Ishi, Jani Even, Norihiro Hagita |
| 2013 | IROS | Using sound reflections to detect moving entities out of the field of view. | Nagasrikanth Kallakuri, Jani Even, Yoichi Morales Saiki, Carlos Toshinori Ishi, Norihiro Hagita |
| 2012 | BIBE | The role of the Lombard reflex in parkinson's disease. | Panikos Heracleous, Jani Even, Carlos Toshinori Ishi, Takahiro Miyashita, Norihiro Hagita, Masaki Kondo, Kyoko Takanohara |
| 2012 | HRI | Generation of nodding, head tilting and eye gazing for human-robot dialogue interaction. | Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita |
| 2012 | ICASSP | Fusion of standard and alternative acoustic sensors for robust automatic speech recognition. | Panikos Heracleous, Jani Even, Carlos Toshinori Ishi, Takahiro Miyashita, Norihiro Hagita |
| 2012 | Interspeech | Evaluation of a formant-based speech-driven lip motion generation. | Carlos Toshinori Ishi, Chaoran Liu, Hiroshi Ishiguro, Norihiro Hagita |
| 2012 | IROS | Combining laser range finders and local steered response power for audio monitoring. | Jani Even, Carlos Toshinori Ishi, Panikos Heracleous, Takahiro Miyashita, Norihiro Hagita |
| 2012 | IROS | Evaluation of formant-based lip motion generation in tele-operated humanoid robots. | Carlos Toshinori Ishi, Chaoran Liu, Hiroshi Ishiguro, Norihiro Hagita |
| 2012 | LREC | Body-conductive acoustic sensors in human-robot communication. | Panikos Heracleous, Carlos Toshinori Ishi, Takahiro Miyashita, Norihiro Hagita |
| 2011 | Interspeech | Range Based Multi Microphone Array Fusion for Speaker Activity Detection in Small Meetings. | Jani Even, Panikos Heracleous, Carlos Toshinori Ishi, Norihiro Hagita |
| 2011 | Interspeech | Improved Acoustic Characterization of Breathy and Whispery Voices. | Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita |
| 2011 | Interspeech | Analysis of Acoustic-Prosodic Features Related to Paralinguistic Information Carried by Interjections in Dialogue Speech. | Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita |
| 2011 | IROS | Multi-modal front-end for speaker activity detection in small meetings. | Jani Even, Panikos Heracleous, Carlos Toshinori Ishi, Norihiro Hagita |
| 2011 | IROS | The effects of microphone array processing on pitch extraction in real noisy environments. | Carlos Toshinori Ishi, Liang Dong, Hiroshi Ishiguro, Norihiro Hagita |
| 2011 | SIGGRAPH | Telenoid: tele-presence android for communication. | Kohei Ogawa, Shuichi Nishio, Kensuke Koda, Koichi Taura, Takashi Minato, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2010 | HRI | Head motions during dialogue speech and nod timing control in humanoid robots. | Carlos Toshinori Ishi, Chaoran Liu, Hiroshi Ishiguro, Norihiro Hagita |
| 2010 | Interspeech | Close speaker cancellation for suppression of non-stationary background noise for hands-free speech interface. | Jani Even, Carlos Toshinori Ishi, Hiroshi Saruwatari, Norihiro Hagita |
| 2010 | IROS | Sound interval detection of multiple sources based on sound directivity. | Carlos Toshinori Ishi, Liang Dong, Hiroshi Ishiguro, Norihiro Hagita |
| 2009 | ACII | How about laughter? Perceived naturalness of two laughing humanoid robots. | Christian Becker-Asano, Toshiyuki Kanda, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2009 | IROS | Evaluation of a MUSIC-based real-time sound localization of multiple sound sources in real noisy environments. | Carlos Toshinori Ishi, Olivier Chatot, Hiroshi Ishiguro, Norihiro Hagita |
| 2008 | HRI | A semi-autonomous communication robot: a field trial at a train station. | Masahiro Shiomi, Daisuke Sakamoto, Takayuki Kanda, Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita |
| 2008 | Interspeech | The meanings carried by interjections in spontaneous speech. | Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita |
| 2007 | Interspeech | Analysis of head motions and speech in spoken dialogue. | Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita |
| 2007 | IROS | Analysis of head motions and speech, and head motion control in an android. | Carlos Toshinori Ishi, Judith Haas, Freerk Pieter Wilbers, Hiroshi Ishiguro, Norihiro Hagita |
| 2007 | IROS | A blendshape model for mapping facial motions to an android. | Freerk Pieter Wilbers, Carlos Toshinori Ishi, Hiroshi Ishiguro |
| 2006 | Interspeech | Analysis of prosodic and linguistic cues of phrase finals for turn-taking and dialog acts. | Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita |
| 2006 | IROS | Evaluation of Prosodic and Voice Quality Features on Automatic Extraction of Paralinguistic Information. | Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita |
| 2005 | Interspeech | Proposal of acoustic measures for automatic detection of vocal fry. | Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita |
| 2004 | Interspeech | A new acoustic measure for aspiration noise detection. | Carlos Toshinori Ishi |
| 2003 | Interspeech | Perceptually-related acoustic-prosodic features of phrase finals in spontaneous speech. | Carlos Toshinori Ishi, Parham Mokhtari, Nick Campbell |
| 2001 | Interspeech | Identification of accent and intonation in sentences for CALL systems. | Carlos Toshinori Ishi, Nobuaki Minematsu, Ryuji Nishide, Keikichi Hirose |
| 2000 | Interspeech | Identification of Japanese double-mora phonemes considering speaking rate for the use in CALL systems. | Carlos Toshinori Ishi, Keikichi Hirose, Nobuaki Minematsu |
| 2000 | Interspeech | The distribution of fillers in lectures in the Japanese language. | Michiko Watanabe, Carlos Toshinori Ishi |
| 1999 | Interspeech | A system for learning the pronunciation of Japanese pitch accent. | Goh Kawai, Carlos Toshinori Ishi |