Multimodal Adaptive Interfaces
Deb Kumar Roy, Alex Pentland · 1998
this paper we describe recent work on multimodal adaptive interfaces which combine automatic speech recognition, computer vision for gesture tracking, and machine learning techniques. Speech is the primary mode of communication between people and should also be used in computer human communication. Gesture usually accompanies speech and provides information which is at times complementary and at times redundant to the information in the speech stream. Depending on the task at hand and the user's preferences, she will use a combination of speech and gesture in different ways to communicate her intent. In this paper we present preliminary results of an interface which lets the user communicate using a combination of speech and diectic (pointing) gestures. Although other efforts have been made to build multimodal interfaces, we present a system which centers around on-line learning to actively acquire communicative primitives from interactions with the user. In this paper we begin in Section 2 by examining some of the problems of designing interfaces which use natural modalities and motivate our approach which centers around enabling the interface to learn from the user. In Section 3 we give an overview of our approach to addressing the issues raised in Section 2. Section 4 describes the multimodal sensory environment we have built for developing our interfaces. This environment includes a vision based hand tracking system and a phonetic speech recognizer. In Section 5 we introduce Toco the Toucan, an animated synthetic character which provides embodiment for the interface. Section 6 details the learning algorithm which allows Toco to learn the acoustic models and meanings of words as the user points to virtual objects and talk about them. In Section 7 we summarize our w...