A Self-Organizing Japanese Word Segmenter using Heuristic Word Identification and Re-estimation
Masaaki Nagata · 1997
We present a self-organized method to build a stochastic Japanese word segmenter/tom a small number of basic words and a large amount of unsegmented triig text. It consists of a word-based statistical language model, an initial estimation procedure, and a re-estimation procedure. Initial word/requencies are estimated by counting all possible longest match strings between the tr.i text and the word list. The initial word list is augmented by identifying words in the training text using a heuristic rule based on character type. The word-based language model is then re-estimated to filter out inappropriate word hypotheses generated by the initial word identification. When the word segmenter is trained on 3.9M character texts and 1719 initial words, its word segmentation accuracy is 86.3% recall and 82.5% precision. We find that the combination of heuristic word identification and re-estimation is so effective that the initial word list need not be large.