A framework for multimodal integration

Anurag Gupta · UNSWorks (University of New South Wales, Sydney, Australia) · 2022

Multimodal interfaces allow users to provide inputs in multiple modalities such as speech and gesture. Multimodal inputs contain cross-modal deictic referential expressions and they can contribute information that are complementary, redundant, logically related, etc. Individual modalities can generate ambiguous interpretations due to erroneous recognition or interpretation of user inputs. The role of multimodal integration is to capture significant multimodal phenomena within a semantic representation formalism, integrate semantically rich and possibly ambiguous interpretations received from multiple modalities to generate multimodal interpretation candidates, and finally determine the correct joint multimodal interpretation. This research was carried out in steps starting with reviewing existing literature to establish the state-of-art techniques, observing behaviour of users with multimodal systems, and then establishing relevant research opportunities. The identified opportunities influenced the design of a standard reference model for multimodal input interpretation that defines the processes required for multimodal integration. The reference model led to the development of an architecture for the Multimodal Input Fusion (MMIF) module and the definition of a Multimodal Feature Structure formalism that represents semantic content and significant multimodal phenomena in interpretations received from input modalities. These were followed by the development and analysis of several algorithms that implement the processes required for multimodal integration. The MMIF module collects interpretations from input modalities and determines the end of turn, i.e., the state where the user expects a system response after providing one or more inputs. The MMIF module disambiguates the collected interpretations to determine appropriate combinations for integration. The combinations of interpretations are integrated to generate multimodal interpretations, which are ranked to determine the correct multimodal interpretation. The MMIF module aggregates and selects possible referents to resolve cross-modal references, and uses the semantic operator based fusion technique to handle semantic-level relationships between the collected interpretations. The MMIF module has been used in several multimodal systems that provide form filling, command & control, and interactive dialogue modes of interaction. The results of evaluation of the MMIF module demonstrate its capability to handle natural multimodal inputs, high accuracy in creating the correct multimodal interpretation, and flexibility to users for providing multimodal inputs.

Read the paper · More papers on PaperTik