Message-driven speech recognition and topic-word extraction

Katsutoshi Ohtsuki, Sadaoki Furui, A. Iwasaki, Naoyuki Sakurai · 1999

This paper proposes a new formulation for speech recognition/understanding systems. In which the posteriori probability of a speaker's message that the speaker intends to address given an observed acoustic sequence is maximized. This is an extension of the current criterion that maximizes the probability of a word sequence. Among the various possible representations, we employ a co-occurrence score of words measured by mutual information as the conditional probability of a word sequence occurring in a given message. The word sequence hypotheses obtained by bigram and trigram language models are rescored using the co-occurrence score. Experimental results show that the word accuracy is improved by this method. Topic-words which represent the content of a speech signal are then extracted from speech recognition results based on the significance score of each word. When five topic-words are extracted for each broadcast-news article, 82.8% of them are correct in average. This paper also proposes a verbalization-dependent language model which is useful for Japanese dictation systems.

Read the paper · More papers on PaperTik