Natural versus Standardized Approaches to Spoken System Design: A Comparison Using the Dual-Task Paradigm
Michael K. Tanenhaus, James F. Allen, Ellen Campana · 2009
The design of spoken dialogue systems/speech interfaces is dominated by two main approaches that are rarely explicitly discussed in literature, which I call the natural and the standardized approach. The natural approach takes as a gold standard human-human interaction, while standardized approach advocates design decisions that reduce computational complexity while providing user with consistent language to facilitate learning and adaptation. There is no need to assume, a priori, that either approach will be more useful to all people in all situations. Comparing two approaches involves pitting powerful human language capacities against powerful human learning abilities, pitting distribution of language a person has produced and understood over course of a lifetime against a smaller and more constrained situation-specific distribution, and possibly even pitting automatic language mechanisms against higher-level perspective-taking mechanisms. Until now, there has been little empirical research directly comparing two approaches in terms of usability, in part due to lack of an evaluation metric that is directly related to ease-of-use, yet fine-grained enough to be used at utterance-level. In this dissertation I extend a classic tool from Cognitive Psychology, dual-task paradigm, to system evaluation in order to address this question. I focus specifically on generation of referring expressions because (1) referring expressions are fundamental to all spoken communication, and (2) the two approaches advocate different methods. I consider role of both discourse context and visual context. With respect to role of discourse context in generation of referring expressions results suggest that systems that take a natural approach consume fewer of user's cognitive resources (specifically visual attention), even after practice. In contrast, with respect to role of visual context in generation of referring expressions results suggest that with practice users can easily adapt to standardized systems that ignore visual context. For dialogue system/speech interface designers implications are that for maximum usability systems should take a natural approach to role of discourse context, but role of visual context is a potential candidate for standardization/streamlining.