An Efficient Unification-Based Multimodal Language Processor for Multimodal Input Fusion
Fang Chen, Yong Sun · IGI Global eBooks · 2010
Multimodal user interaction technology aims at building natural and intuitive interfaces allowing a user to interact with computers in a way similar to human-to-human communication, for example, through speech and gestures. As a critical component in a multimodal user interface, multimodal input fusion explores ways to effectively derive the combined semantic interpretation of user inputs through multiple modalities. Based on state–of-the-art review on multimodal input fusion approaches, this chapter presents a novel approach to multimodal input fusion based on speech and gesture; or speech and eye tracking. It can also be applied for other input modalities and extended to more than two modalities. It is the first time that a powerful combinational categorical grammar is adopted in multimodal input fusion. The effectiveness of the approach has been validated through user experiments, which indicated a low polynomial computational complexity while parsing versatile multimodal input patterns. It is very useful for mobile context. Future trends in multimodal input fusion will be discussed at the end of this chapter.