| 2004 | Private speech during multimodal human-computer interaction. | Rebecca Lunsford |
| 2004 | Multimodal interface design for multimodal meeting content retrieval. | Agnes Lisowska |
| 2004 | Two-way eye contact between humans and robots. | Yoshinori Kuno, Arihiro Sakurai, Dai Miyauchi, Akio Nakamura |
| 2004 | Multimodal interaction under exerted conditions in a natural field setting. | Sanjeev Kumar, Philip R. Cohen, Rachel Coulston |
| 2004 | Towards integrated microplanning of language and iconic gesture for multimodal output. | Stefan Kopp, Paul Tepper, Justine Cassell |
| 2004 | Exploiting prosodic structuring of coverbal gesticulation. | Sanshzar Kettebekov |
| 2004 | Identifying the addressee in human-human-robot interactions based on head pose and speech. | Michael Katzenmaier, Rainer Stiefelhagen, Tanja Schultz |
| 2004 | A multimodal learning interface for sketch, speak and point creation of a schedule chart. | Edward C. Kaiser, David Demirdjian, Alexander Gruenstein, Xiaoguang Li, John Niekrasz, Matt Wesson, Sanjeev Kumar |
| 2004 | Elvis: situated speech and gesture understanding for a robotic chandelier. | Joshua Juster, Deb Roy |
| 2004 | Multilayer architecture in sign language recognition system. | Feng Jiang, Hongxun Yao, Guilin Yao |
| 2004 | Implementation and evaluation of a constraint-based multimodal fusion system for speech and 3D pointing gestures. | Hartwig Holzapfel, Kai Nickel, Rainer Stiefelhagen |
| 2004 | Using spatial warning signals to capture a driver's visual attention. | Cristy Ho |
| 2004 | Multimodal interaction in an augmented reality scenario. | Gunther Heidemann, Ingo Bax, Holger Bekel |
| 2004 | A segment-based audio-visual speech recognizer: data collection, development, and initial experiments. | Timothy J. Hazen, Kate Saenko, Chia-Hao La, James R. Glass |
| 2004 | Multimodal model integration for sentence unit detection. | Mary P. Harper, Elizabeth Shriberg |
| 2004 | Agent and library augmented shared knowledge areas (ALASKA). | Eric R. Hamilton |
| 2004 | An evaluation of virtual human technology in informational kiosks. | Curry I. Guinn, Robert C. Hubal |
| 2004 | M/ORIS: a medical/operating room interaction system. | Sbastien Grange, Terrence Fong, Charles Baur |
| 2004 | Software infrastructure for multi-modal virtual environments. | Brian F. Goldiez, Glenn A. Martin, Jason Daly, Donald Washburn, Todd Lazarus |
| 2004 | User walkthrough of multimodal access to multidimensional databases. | Myra P. van Esch-Bussemakers, Anita H. M. Cremers |
| 2004 | Visual and linguistic information in gesture classification. | Jacob Eisenstein, Randall Davis |
| 2004 | Gestural cues for speech understanding. | Jacob Eisenstein |
| 2004 | Support for input adaptability in the ICON toolkit. | Pierre Dragicevic, Jean-Daniel Fekete |
| 2004 | Real-time audio-visual tracking for meeting analysis. | David Demirdjian, Kevin W. Wilson, Michael Siracusa, Trevor Darrell |
| 2004 | Command and control resource performance predictor(C | Joseph M. Dalton, Ali Ahmad, Kay M. Stanney |