Multimodal Interfaces – A Generic Design Approach

Noëlle Carbonell · 2005

Integrating new input-output modalities, such as speech, gaze, gestures, haptics, etc., in user interfaces is currently considered as a significant potential contribution to implementing the concept of Universal Access (UA) in the Information Society (see, for example, Oviatt, 2003). UA in this context means providing everybody, including disabled users, with easy humancomputer interaction in any context of use, and especially in mobile contexts. However, the cost of developing an appropriate specific multimodal user interface for each interactive software is prohibitive. A generic design methodology, along with generic reusable components, are needed to master the complexity of the design and development of interfaces that allow flexible use of alternative modalities, in meaningful combinations, according to the constraints in the interaction environment or the user's motor and perceptual capabilities. We present a design approach meant to facilitate the development of generic multimodal user interfaces, based on best practice in software and user interface design and architecture. 1. Problem Being Addressed At present, specific off the shelf components are available that can process and interpret data from a wide range of input devices reliably. However, these components are monomodal, in the sense that they are dedicated to a specific medium and modality. There is not yet, outside research laboratory prototypes, any software platform capable of interpreting multimodal input data. Similarly, software on the market is available for the generation of monomodal output messages conveyed through various media, whereas the generation of multimodal presentations is not yet supported. In the next section, we present a design approach that makes it possible to: • Interpret users' multimodal commands or manipulations/actions using partial monomodal interpretations, each interpretation being elaborated by a dedicated component which processes the specific input data stream transmitted through one of the available media. This treatment may be viewed as a fusion process of events or data. 210 Noёlle Carbonell • Match these global interpretations with appropriate functions in the current application software, or translate them into appropriate commands (i.e., execution calls of the appropriate functions in the kernel of the considered software). • As regards system multimodal outputs, break up the information content of system messages into chunks, assign the resulting data chunks to appropriate modalities, and input each of them into the relevant monomodal generation/presentation component. This treatment is often viewed as a data fission process in contrast with multimodal input fusion. In most current applications, stereotyped system messages only need to be implemented in order to achieve efficient interaction, so that simple techniques can be applied to generate appropriate multimodal system messages. Modality selection can be easily performed using available ergonomic criteria. As regards accessibility, the World Wide Web Consortium's (W3C) Web Access Initiative (WAI) has designed accessibility guidelines (http://www.w3.org/WAI/) which have been implemented in software tools such as Bobby1. Efficient generation software components also exist for many output modalities (except for haptics, a which still needs further research studies), namely, speech, sound, graphics and text. The fusion of multimodal inputs and their global interpretation in terms of executable function calls are, on the other hand, much more complex, due mainly to the relative complexity of users' utterances and the limitations of current monomodal interpreters, especially concerning modalities such as speech, gestures and haptics. The approach presented in the remainder of this Chapter is meant to help designers to overcome these difficulties in contexts of use where the semantics of the information exchanges between user and software amounts only to the expressive power of Direct Manipulation (Shneiderman, 1993). As this Chapter is focused on software architecture issues, it does not include some of the sections that are present in the other Chapters, namely: • The 'Procedures for using the device' section, since numerous software design methods exist, and most companies have evolved their own design and development methodology, practice and standards. In addition, software design issues have motivated the publication of numerous manuals and best practice case studies; see, for instance, (Bass et al., 1998) for design issues, and (Clements et al., 2001) for evaluation methods. • 'Outcomes' or rather the advantages of the software architecture proposed for multimodal user interfaces are briefly discussed in the relevant section of this Chapter. • The 'Assumptions' section is not relevant to this Chapter: as the proposed architecture framework aims at genericity, no restriction should be placed on its applicability. 1 Bobby (http://bobby.watchfire.com/), a Watchfire Corporation product, is meant to help Web application designers to comply with these guidelines by exposing accessibility problems in Web pages and suggesting solutions to repair them. Chapter 17 Multimodal Interfaces – A Generic Design Approach 211 2. Device / Technique(s) Used After providing brief definitions of what we mean by modality and multimodality, we shall present a generic overall software architecture for multimodal user interfaces. This generic architecture is based on a five-layer software model, each layer being designed according to component programming principles. Intercomponent information exchanges implement the W3C SOAP message exchange paradigm2. It is meant to facilitate multimodal input and output processing. In particular, the interpretation of users' multimodal utterances simply consists of matching them with the appropriate functions in the considered specific application software, in order to achieve robust human-computer interaction while limiting development complexity and cost. As it is largely application-independent, generic interpreters and generators on the market can be used for processing monomodal inputs and outputs.

Read the paper · More papers on PaperTik