| 2012 | A hierarchical approach to continuous gesture analysis for natural multi-modal interaction. | Ying Yin |
| 2012 | Simple multi-party video conversation system focused on participant eye gaze: "Ptolemaeus" provides participants with smooth turn-taking. | Saori Yamamoto, Nazomu Teraya, Yumika Nakamura, Narumi Watanabe, Yande Lin, Mayumi Bono, Yugo Takeuchi |
| 2012 | Development of the 2012 SJTU HVR system. | Hainan Xu, Yuchen Fan, Kai Yu |
| 2012 | Multimodal detection of salient behaviors of approach-avoidance in dyadic interactions. | Bo Xiao, Panayiotis G. Georgiou, Brian R. Baucom, Shrikanth S. Narayanan |
| 2012 | Multimodal learning analytics: enabling the future of learning through multimodal data analysis and interfaces. | Marcelo Worsley |
| 2012 | Design space for finger gestures with hand-held tablets. | Katrin Wolf |
| 2012 | Designing multimodal reminders for the home: pairing content with presentation. | Julie R. Williamson, Marilyn Rose McGee-Lennon, Stephen A. Brewster |
| 2012 | Electroencephalographic detection of visual saliency of motion towards a practical brain-computer interface for video analysis. | Matthew Weiden, Deepak Khosla, Matthew Keegan |
| 2012 | Let's have dinner together: evaluate the mediated co-dining experience. | Jun Wei, Adrian David Cheok, Ryohei Nakatsu |
| 2012 | A framework of personal assistant for computer users by analyzing video stream. | Zixuan Wang, Jinyun Yan, Hamid K. Aghajan |
| 2012 | Improving mandarin predictive text input by augmenting pinyin initials with speech and tonal information. | Guangsen Wang, Bo Li, Shilin Liu, Xuancong Wang, Xiaoxuan Wang, Khe Chai Sim |
| 2012 | Hard lessons learned: mobile eye-tracking in cockpits. | Hana Vrzakova, Roman Bednarik |
| 2012 | A portable audio/video recorder for longitudinal study of child development. | Soroush Vosoughi, Matthew S. Goodwin, Bill Washabaugh, Deb Roy |
| 2012 | Learning speaker, addressee and overlap detection models from multimodal streams. | Oriol Vinyals, Dan Bohus, Rich Caruana |
| 2012 | Pixene: creating memories while sharing photos. | Ramadevi Vennelakanti, Sriganesh Madhvanath, Anbumani Subramanian, Ajith Sowndararajan, Arun David, Prasenjit Dey |
| 2012 | Gestures as point clouds: a $P recognizer for user interface prototypes. | Radu-Daniel Vatavu, Lisa Anthony, Jacob O. Wobbrock |
| 2012 | NeuroDialog: an EEG-enabled spoken dialog interface. | Seshadri Sridharan, Yun-Nung Chen, Kai-min Chang, Alexander I. Rudnicky |
| 2012 | Multimodal human behavior analysis: learning correlation and interaction across modalities. | Yale Song, Louis-Philippe Morency, Randall Davis |
| 2012 | A multimodal fuzzy inference system using a continuous facial expression representation for emotion detection. | Catherine Soladi, Hanan Salam, Catherine Pelachaud, Nicolas Stoiber, Renaud Sguier |
| 2012 | IrisTK: a statechart-based toolkit for multi-party face-to-face interaction. | Gabriel Skantze, Samer Al Moubayed |
| 2012 | ICMI'12 grand challenge: haptic voice recognition. | Khe Chai Sim, Shengdong Zhao, Kai Yu, Hank Liao |
| 2012 | Speak-as-you-swipe (SAYS): a multimodal interface combining speech and gesture keyboard synchronously for continuous mobile text entry. | Khe Chai Sim |
| 2012 | Investigating the midline effect for visual focus of attention recognition. | Samira Sheikhi, Jean-Marc Odobez |
| 2012 | Changes in verbal and nonverbal conversational behavior in long-term interaction. | Daniel Schulman, Timothy W. Bickmore |
| 2012 | AVEC 2012: the continuous audio/visual emotion challenge. | Bjrn W. Schuller, Michel F. Valstar, Florian Eyben, Roddy Cowie, Maja Pantic |