| 2025 | ICASSP | Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings. | Jason Clarke, Yoshihiko Gotoh, Stefan Goetze |
| 2023 | ASRU | Improving Audiovisual Active Speaker Detection in Egocentric Recordings with the Data-Efficient Image Transformer. | Jason Clarke, Yoshihiko Gotoh, Stefan Goetze |
| 2023 | CHI | Exploration of verbal descriptions and dynamic indoors environments for people with sight loss. | Abdulaziz Alrashidi, Peter Cudd, Charith Abhayaratne, Yoshihiko Gotoh |
| 2019 | ICASSP | 3D Visual Speech Animation Using 2D Videos. | Rabab Algadhy, Yoshihiko Gotoh, Steve Maddock |
| 2018 | BMVC | Graph-based Correlated Topic Model for Motion Patterns Analysis in Crowded Scenes from Tracklets. | Manal Alghamdi, Yoshihiko Gotoh |
| 2018 | WACV | Graph-Based Correlated Topic Model for Trajectory Clustering in Crowded Videos. | Manal Al Ghamdi, Yoshihiko Gotoh |
| 2017 | INLG | Natural Language Descriptions for Human Activities in Video Streams. | Nouf Al Harbi, Yoshihiko Gotoh |
| 2016 | ACL | Natural Language Descriptions of Human Activities Scenes: Corpus Generation and Analysis. | Nouf Al Harbi, Yoshihiko Gotoh |
| 2014 | ECIR | Video Clip Retrieval by Graph Matching. | Manal Al Ghamdi, Yoshihiko Gotoh |
| 2014 | ICASSP | Alignment of nearly-repetitive contents in a video stream with manifold embedding. | Manal Al Ghamdi, Yoshihiko Gotoh |
| 2014 | ICISP | Manifold Matching with Application to Instance Search Based on Video Queries. | Manal Al Ghamdi, Yoshihiko Gotoh |
| 2013 | CAIP | Spatio-temporal Manifold Embedding for Nearly-Repetitive Contents in a Video Stream. | Manal Al Ghamdi, Yoshihiko Gotoh |
| 2013 | CAIP | Spatio-temporal Human Body Segmentation from Video Stream. | Nouf Al Harbi, Yoshihiko Gotoh |
| 2012 | ECCV | Spatio-temporal Video Representation with Locality-Constrained Linear Coding. | Manal Al Ghamdi, Nouf Al Harbi, Yoshihiko Gotoh |
| 2012 | ECCV | Spatio-temporal SIFT and Its Application to Human Action Classification. | Manal Al Ghamdi, Lei Zhang, Yoshihiko Gotoh |
| 2012 | ICIP | Generating coherent natural language annotations for video streams. | Muhammad Usman Ghani Khan, Lei Zhang, Yoshihiko Gotoh |
| 2010 | ICASSP | Nearly-repetitive video synchronisation using nonlinear manifold embedding. | Siripinyo Chantamunee, Yoshihiko Gotoh |
| 2007 | Interspeech | Relative evaluation of informativeness in machine generated summaries. | BalaKrishna Kolluru, Yoshihiko Gotoh |
| 2007 | Interspeech | Speaker role based structural classification of broadcast news stories. | BalaKrishna Kolluru, Yoshihiko Gotoh |
| 2005 | ACL | On the Subjectivity of Human Authored Summaries. | BalaKrishna Kolluru, Yoshihiko Gotoh |
| 2005 | ICASSP | Maximum entropy segmentation of broadcast news. | Heidi Christensen, BalaKrishna Kolluru, Yoshihiko Gotoh, Steve Renals |
| 2005 | Interspeech | Multi-stage compaction approach to broadcast news summarisation. | BalaKrishna Kolluru, Heidi Christensen, Yoshihiko Gotoh |
| 2004 | ECIR | From Text Summarisation to Style-Specific Summarisation for Broadcast News. | Heidi Christensen, BalaKrishna Kolluru, Yoshihiko Gotoh, Steve Renals |
| 2000 | ICASSP | Variable word rate N-grams. | Yoshihiko Gotoh, Steve Renals |
| 1999 | ICASSP | Named entity tagged language models. | Yoshihiko Gotoh, Steve Renals, Gethin Williams |
| 1999 | Interspeech | Integrated transcription and identification of named entities in broadcast speech. | Steve Renals, Yoshihiko Gotoh |
| 1997 | Interspeech | Document space models using latent semantic analysis. | Yoshihiko Gotoh, Steve Renals |
| 1996 | ICASSP | Microphone-array speech recognition via incremental map training. | John E. Adcock, Yoshihiko Gotoh, Daniel J. Mashao, Harvey F. Silverman |
| 1996 | ICASSP | Incremental ML estimation of HMM parameters for efficient training. | Yoshihiko Gotoh, Harvey F. Silverman |
| 1994 | ICASSP | Using MAP estimated parameters to improve HMM speech recognition performance. | Yoshihiko Gotoh, Michael M. Hochberg, Harvey F. Silverman |