Multimodal Interfaces: A Study on Speech-Hand Gesture Recognition
Jude Joseph Lamug Martinez, Sindy Senorita Dewanti · 2019 International Conference on Information and Communications Technology (ICOIACT) · 2019
One common unimodal human-computer interface is via a formal and discrete behavior on physical or virtual button clicks. That method may not be convenient for some people with knowledge or physical limitations. The raise for more pervasive, subtly blending computers in daily tasks may also find the method to be unnatural and inconvenient. Combination of modalities, or multimodal interaction is suggested to have advantages over unimodal input. Synergic multiple input and output modal is more expressive and natural. Specifically, studies have suggested that a combination of speech and gesture is preferred as being more effective and natural than speech or gesture alone. There are many successful approaches in providing robust gesture-only and speech-only interfaces, however, under extreme environments, unimodal approaches have user experience issues. To address the issue, the authors seek to explore and conduct a preliminary study and evaluation of speech-hand gesture recognition using Leap Motion and Windows Speech Recognition API. A conceptual framework would be developed as a result of this study.