N-gram language modeling of Japanese using bunsetsu boundaries
Sungyup Chung, Keikichi Hirose, Nobuaki Minematsu · 2004
A new scheme of N-gram language modeling was pro-posed for Japanese, where word N-grams were calculated separately for the two cases: crossing and not crossing bunsetsu boundaries. Here, bunsetsu is a basic gram-matical (and pronunciation) unit of Japanese. A similar scheme using accent phrase boundaries instead of bun-setsu boundaries has already been proposed by the au-thors with a certain success, but it suffered from the train-ing data shortage, because assignment of accent phrase boundaries requires a speech corpus. In contrast, bun-setsu boundaries can be detected automatically from a written text with a rather high accuracy using a parser. It was shown from the experiment that a perplexity re-duction was possible by estimating bunsetsu boundaries from the history longer than N-1 words in the case of N-gram modeling and by selecting one from two types of models (crossing and not crossing bunsetsu bound-aries) according to the estimation. When 1 or 3 years of Mainichi Newspaper corpus was used for the training of tri-grams, the proposed scheme could reduce the perplex-ity by around 8 % from the baseline modeling (without separation). The proposed language modeling was ap-plied to a continuous speech recognition, and it showed that an improvement in word recognition rate was possi-ble especially when the training corpus was small (1 year of newspaper). 1.