| 2006 | Speaker localization for microphone array-based ASR: the effects of accuracy on overlapping speech. | Hari Krishna Maganti, Daniel Gatica-Perez |
| 2006 | Gaze-X: adaptive affective multimodal interface for single-user office scenarios. | Ludo Maat, Maja Pantic |
| 2006 | Toward open-microphone engagement for multiparty interactions. | Rebecca Lunsford, Sharon L. Oviatt, Alexander M. Arthur |
| 2006 | Human perception of intended addressee during computer-assisted meetings. | Rebecca Lunsford, Sharon L. Oviatt |
| 2006 | Word graph based speech rcognition error correction by handwriting input. | Peng Liu, Frank K. Soong |
| 2006 | Evaluating usability based on multimodal information: an empirical study. | Tao Lin, Atsumi Imamiya |
| 2006 | A fast and robust 3D head pose and gaze estimation system. | Koichi Kinoshita, Yong Ma, Shihong Lao, Masato Kawade |
| 2006 | Short message dictation on Symbian series 60 mobile phones. | E. Karpov, Imre Kiss, Jussi Leppnen, Jesper . Olsen, Daniela Oria, S. Sivadas, Jilei Tian |
| 2006 | Using redundant speech and handwriting for learning new vocabulary and understanding abbreviations. | Edward C. Kaiser |
| 2006 | Human-Robot dialogue for joint construction tasks. | Mary Ellen Foster, Tomas By, Markus Rickert, Alois C. Knoll |
| 2006 | The NIST smart data flow system II multimodal data transport infrastructure. | Antoine Fillinger, Stphane Degr, Imad Hamchi, Vincent Stanford |
| 2006 | Haptic phonemes: basic building blocks of haptic communication. | Mario J. Enriquez, Karon E. MacLean, Christian Chita |
| 2006 | A 'need to know' system for group classification. | Wen Dong, Jonathan Gips, Alex Pentland |
| 2006 | MyConnector: analysis of context cues to predict human availability for communication. | Maria Danninger, Tobias Kluge, Rainer Stiefelhagen |
| 2006 | Foundations of human computing: facial expression and emotion. | Jeffrey F. Cohn |
| 2006 | Mixing virtual and actual. | Herbert H. Clark |
| 2006 | Co-Adaptation of audio-visual speech and gesture classifiers. | C. Mario Christoudias, Kate Saenko, Louis-Philippe Morency, Trevor Darrell |
| 2006 | Modeling naturalistic affective states via facial and vocal expressions recognition. | George Caridakis, Lori Malatesta, Loc Kessous, Noam Amir, Amaryllis Raouzaiou, Kostas Karpouzis |
| 2006 | Comparing the effects of visual-auditory and visual-tactile feedback on user performance: a meta-analysis. | Jennifer L. Burke, Matthew S. Prewett, Ashley A. Gray, Liuquin Yang, Frederick R. B. Stilson, Michael D. Coovert, Linda R. Elliott, Elizabeth S. Redden |
| 2006 | A contextual multimodal integrator. | Pter Pl Boda |
| 2006 | Computing human faces for human viewers: automated animation in photographs and paintings. | Volker Blanz |
| 2006 | CarDialer: multi-modal in-vehicle cellphone control application. | Vladimr Bergl, Martin Cmejrek, Martin Fanta, Martin Labsk, Ladislav Serdi, Jan Sediv, Lubos Ures |
| 2006 | Collaborative multimodal photo annotation over digital paper. | Paulo Barthelmess, Edward C. Kaiser, Xiao Huang, David McGee, Philip R. Cohen |
| 2006 | Collaborative multimodal photo annotation over digital paper. | Paulo Barthelmess, Edward C. Kaiser, Xiao Huang, David McGee, Philip R. Cohen |
| 2006 | Prototyping novel collaborative multimodal systems: simulation, data collection and analysis tools for the next decade. | Alexander M. Arthur, Rebecca Lunsford, Matt Wesson, Sharon L. Oviatt |