Multiple utterance prediction based on a tree structure of dialogue states

Hiromasa Terashi, Masafumi Nishida, Yasuo Horiuchi, Akira Ichikawa · The Journal of the Acoustical Society of America · 2006

A problem of spoken dialogue systems is that recognition errors increase when the recognition-word vocabulary is large. To solve the problem, we propose a method to predict a user’s utterances by recognizing prediction sentences of each dialogue state. The method can judge whether the user’s utterance is within the range of prediction using decoders based on a prediction sentence and a large word vocabulary. However, the conventional method performs processing to lead in a prediction frequently when a topic changes because even an utterance within a domain is treated as an utterance of the prediction outside. In this study, we propose a new method using multiple predictions by performing recognition based on all prediction sentences in a domain. The proposed method produces a tree structure for dialogue states in a domain and regards child nodes of an arbitrary node as within the range of prediction. We conducted experiments for 451 utterances of 15 dialogues by ten persons using the university guide system. As a result, recognition accuracy of prediction sentences was 97.7% and the judgment accuracy of utterance predictions was 88.2%; high performance was obtained using the proposed method. Results demonstrate that it is possible to perform flexible dialogue control.

Read the paper · More papers on PaperTik