Method for predicting utterances using a layered plan recognition model—disambiguation of speech recognition candidates by focusing on speaker intentions

Takayuki Yamaoka, Hitoshi Iida · Systems and Computers in Japan · 1994

Abstract A context‐sensitive method that predicts a subsequent utterance to reduce ambiguity in speech recognition candidates for use in a spoken‐language processing system is described. Current speech recognition technology cannot accomplish perfect speech recognition, and in many cases ambiguities remain in the results. On the other hand, linguistic processing techniques assume error‐free input. Therefore the effectiveness of this processing is reduced when there are multiple candidates in the input. In particular, in spite of the importance to dialogue translation of the degree of ambiguity in the sentence‐final expressions that indicate the speaker's intention and colloquial fragmentary expressions that are characteristic of spoken language, conventional techniques do not address this problem. This paper describes first a method for dialogue understanding by means of a layered plan recognition method which uses knowledge on the conduct of a dialogue as well as knowledge concerning the topic, and the use of stacks to manage the states of understand‐ing that this method produces. Then a method is presented for obtaining abstract contextual information concerning the next utterance by referring to the stack contents. Next, a method of applying that contextual information in the selection of speech recognition candidates is described. Finally, pragmatic knowledge concerning utterance intention is specified, and the effectiveness of these methods in reducing the ambiguity of expressions concerning utterance intention is demonstrated by experiment.

Read the paper · More papers on PaperTik