MULTIMODAL INTERFACES THAT PROCESS WHAT COMES NATURALLY

Sharon L. Oviatt, Philip R. Cohen · 2000

During multimodal communication, we speak, shift eye gaze, gesture, and move in a powerful flow of communication that bears little resemblance to the discrete key-board and mouse clicks entered sequentially with a graphical user interface (GUI). A pro-found shift is now occurring toward embracing users ’ natural behavior as the center of the human-computer interface. Multimodal interfaces are being developed that permit our highly skilled and coordinated communicative behavior to control system interactions in a more transparent experience than ever before. Our voice, hands, and entire body, once augmented by sensors such as microphones and cameras, are becoming the ultimate transparent and mobile multimodal input devices. The area of multimodal systems has expanded rapidly during the past five years. Since Bolt’s [1] original “Put That There ” concept demonstration, which processed speech and manual pointing during object manipulation, significant achievements have been made Using our highly skilled and coordinated communication patterns to control computers in a more transparent interface experience. in developing more general multimodal systems. State-of-the-art multimodal speech and gesture systems now process complex gestural input other than pointing, and new systems have been extended to process different mode combinations—the most noteworthy being speech and pen input [9], and speech and lip movements [10]. As a foundation for advancing new multimodal systems, proactive empirical work has generated predictive information on human-computer multimodal interaction, which is being used to

Read the paper · More papers on PaperTik