The limits of speech recognition
Ben Shneiderman · Communications of the ACM · 2000
Spoken language is effective for human-human interaction but often has severe limitations when applied to human-computer interaction. To improve speech recognition applications, designers must understand acoustic memory and prosody, which is the focus of this article. By understanding the cognitive processes surrounding human "acoustic memory" and processing, interface designers may be able to integrate speech more effectively and guide users more successfully. The interfering effects of acoustic processing are a limiting factor for designers of speech recognition, but the role of emotive prosody raises further concerns. The emotive aspects of prosody are potent for human-human interaction but may be disruptive for human-computer interaction. The syntactic aspects of prosody, such as rising tone for questions, are important for a system's recognition and generation of sentences. Better human multitasking models, and an understanding of how human-computer interaction is different from human-human interaction would be helpful in this regard.