Using cognitive models to understand multimodal processes: the case for speech and gesture production
Stefan Kopp, Kirsten von Bergmann · ACM eBooks · 2017
Multimodal behavior has been studied for a long time and in many fields, e.g., in psychology, linguistics, communication studies, education, and ergonomics. One of the main motivations has been to allow humans to use technical systems intuitively, in a way that resembles and fosters human users' natural way of interacting and thinking [Oviatt 2013]. This has sparked early work on multimodal human-computer interfaces, including recent approaches to recognize communicative behavior and even subtle multimodal cues by computer systems. Those approaches, for the most part, rest on machine learning techniques applied to large sets of behavioral data. As datasets grow larger in size and coverage, and computational power increases, suitable data-driven techniques are able to detect correlational behavior patterns that support answering questions like which feature( s) to take into account or how to recognize them in specific contexts. However, natural multimodal interaction in humans entails a plethora of behavioral variations and intricacies (e.g., when to act unimodally vs. multimodally, with which specific behaviors or multi-level coordination between them). Possible underlying patterns are hard to detect, even in large datasets, and often such variations are attributed to context-dependencies or individual differences. How they come about is still hard to explain at the behavioral level.