Towards usable multimodal command languages: definition and ergonomic assessment of constraints on users' spontaneous speech and gestures

Sandrine Robbe, Noëlle Carbonell, Claude Valot · 1997

Within the framework of a prospective ergonomic approach, we simulated two multimodal user interfaces, in order to study the usability of constrained vs spontaneous speech in a multimodal environment. The first experiment, which served as a reference, gave subjects the opportunity to use speech and gestures freely, while subjects in the second experiment had to comply with multimodal constraints. We first describe the experimental setup and the approach we adopted for designing the artificial command language used in the second experiment. We then present the results of our analysis of the subjects’ utterances and gestures, laying emphasis on their implementation of linguistic constraints. The conclusions of the empirical assessment of the usability of this multimodal command language built from a restricted subset of natural language and simple designation gestures is associated with recommendations which may prove useful for improving the usability of oral humancomputer interaction in a multimodal environment. 1 CONTEXT AND MOTIVATION Thanks to recent research advances, speech recognizers are now capable of processing large subsets of natural language (NL) accurately. Nevertheless, spontaneous speech cannot yet be considered as a reliable substitute for artificial query/command languages, menu-driven Human-Computer Interaction (HCI) or direct manipulation, since the interpretation of linguistic reference phrases, especially anaphoric and spatiotemporal phrases, raises still numerous research issues. An attractive solution for achieving robust quasinatural HCI in the near future is to design multimodal languages that allow users to combine oral commands from a restricted subset of NL, with pointing gestures. Such languages are tractable, since multimodal spatial reference phrases (i.e. deictics associated with pointing gestures) can be reliably processed by present NL and gesture interpreters [4]. They may then supersede present forms of HCI, provided that their utility and usability (cf. [3]) is demonstrated beyond doubt. We are currently investigating, within a prospective ergonomic research framework, utility and usability issues raised by the implementation of such languages. The main goal of the comparative empirical study reported here is to assess the effects of realistic expression constraints on the behaviours and attitudes of potential users of forthcoming multimodal interfaces integrating speech and gestures. This study addresses the following major issue: Is it possible to define expression constraints which can both restrict users' spontaneous speech and gestures to a tractable sub-language without interfering with their activity, and be mastered easily in the course of interaction? And if so, how could such constraints be determined? To answer these questions, we considered two experimental situations: in the first one, which served as a reference, subjects could use speech and gestures freely, whereas in the second one they had to comply with multimodal expression constraints. While constrained oral HCI has motivated numerous experimental and empirical studies (cf. for instance, [2], [6]), less attention has been paid to the usability of speech in a multimodal HCI environment. In addition, no comparative study of constrained vs unconstrained oral interaction has been published thus far, at least to our knowledge, and the method used for defining expression constraints is original. 2 EXPERIMENTAL SETUP Two groups, of eight subjects each, interacted during three weekly sessions (of about half an hour per subject each) with two different multimodal interfaces whose functionalities were simulated thanks to the Wizard of Oz technique (WOZ). Both groups carried out identical design tasks; but subjects in the reference group (SP) could use speech and/or 2D gestures (on a touch panel) spontaneously, while subjects in the experimental group (CS) had to comply with expression constraints. 2.1 Expression constraints In order to obtain a tractable multimodal artificial language that would impose minimum constraints on the spontaneous expression of CS subjects, we selected an appropriate subset of the overall set of utterances and gestures used by SP subjects. Two categories of elementary 2D gestures were allowed: pointing gestures, and simulation gestures (akin to mouse drags) for miming translations and rotations of icons on the screen. Ambiguous gestures were eliminated from the simple gestural « vocabulary » used by SP subjects. The verbal component of the language consists in a restricted subset of NL with the following properties. Its syntax can be described by a CF grammar (1) defined on a hundred word vocabulary. Its expressiveness is equivalent to the union of the semantic interpretations of the oral commands issued by SP subjects over the three sessions. Synonymy and polysemy are excluded. CS subjects were given a written description of this multimodal artificial language (2), and the experimenter assisted them while they performed a small set of predefined commands; this initial training stage lasted less than 10 minutes on average. On the other hand, SP subjects had no supervised initial training, but they could, before processing the first scenario, explore the capabilities of the interface in the presence of the experimenter who just answered their questions. 2.2 Application domain and subjects' tasks Subjects had to design or modify furniture arrangements according to instructions specified in scenarios of increasing complexity. Initial furniture layouts were displayed on the screen in the form of 2D plans. 2.3 Implementation of the Wizard of Oz technique Two human operators, hidden from the subjects, simulated the functionalities of both multimodal interfaces. One of them interpreted incoming commands and activated the corresponding software functions which were displayed on the subject's and the wizards' screens; the other one interacted verbally with subjects using a set of fifty or so pre-recorded oral messages. In addition, the CS setup included a commercial continuous speech monospeaker recognizer (Datavox) which the wizards used for interpreting subjects' utterances. 2.4 Recordings and transcripts Subjects were videotaped throughout both experiments. Written descriptions of the recordings comprise orthographic transcripts of verbal exchanges and coarse standardized descriptions of subjects' gestures and system actions; subjects' speech and gestures were further characterized with a view to assessing the extent to which they succeeded in mastering the given set of expression constraints. (1) static branching factor 5.5, dynamic branching factor 2.6 (2) We limited the description of the oral component of the language to the listing of its vocabulary and the presentation of a few commands instances; these instances were chosen so as to illustrate the structural flexibility and expressive power of the language as well as its syntactic and semantic limitations. 3 GLOBAL RESULTS Results presented in sections 3 and 4 bear upon the first session only, since our present main objective (cf. section 1) is to assess the usability of artificial multimodal command languages designed according to the method presented in paragraph 2.1. 3.1 Expression constraints Subjects in the CS group complied easily with gestural constraints: we picked out only three « incorrect » gestures in the transcripts. This result is not surprising since subjects were allowed to use a small number of simple intuitive gestures. On the other hand, all subjects resorted to words outside the vocabulary, and six out of eight used NL structures outside the scope of the language; three subjects only resorted to incorrect (with respect to NL) syntactic structures, while five used NL words or phrases inappropriately. Subjects' reactions to linguistic and enunciation constraints are detailed in section 4. 3.2 Comparison between the CS and SP groups An inter-group comparison suggests that subjects in the CS group benefited from the linguistic constraints with which they had to comply. Hesitations and grammatical errors are significantly (3) less frequent in their oral statements than in those from SP subjects: 13% vs 53% and 1.9% vs 23.5% respectively (4). On the other hand, inter-group differences concerning the use of modalities are not statistically significant by reason of marked inter-individual variations (SCS=67 W(8;8)=]49;87[; p<0.05). Detailed results of this comparative study are presented in [5]. 4 SPEECH CONSTRAINTS We describe here how CS subjects reacted to the linguistic and enunciation constraints they had to comply with. As behaviours and strategies vary greatly from one subject to another, we analyzed the CS transcripts subject by subject, with a view to defining accurate user profiles. Results are summarized in Table 1 and 2. 4.1 Inter-individual variations Inter-individual differences affect many aspects of the verbal expression of subjects as shown in Table 1. First, the total number (NBT) of oral (and multimodal) statements per subject ranges from 4 to 118 (cf. line 4). Recognition rates (NRC/NBC) of the first formulations of correct commands (i.e. commands belonging to the (3) Wilcoxon test: SCS=43 W(8;8)=]49;87[; p<0.05. We applied this test, since there was a significant difference between the variances for the CS and SP groups. (4) Percentages represent numbers of relevant tokens normalized by the total number of oral and multimodal statements per group. CF language) vary also greatly from one subject to another: from 21% of failures up to 57%. S1 S2 S3 S4 S5 S6 S7 S8 NBC 33 30 7 1 7 4 40 0 NRC 7 14 4 0 3 1 9 0 NBE 13 16 3 0 2 6 4 4 NBT 61 118 18 4 20 20 57 12 NTN 20 32 8 1 11 11 14 6 NRR 10 64 6 1 4 6 10 5

Read the paper · More papers on PaperTik