A generalized algorithm for Japanese morphological analysis and a comparative evaluation of some heuristics

Toru Hisamitsu, Yoshihiko Nitta · Systems and Computers in Japan · 1995

Abstract In ordinary written Japanese, words are not separated by spaces. Therefore morphological analysis involves segmenting and tagging sentences. Since each sentence has a huge number of possible tagged segmentations, various criteria have been proposed for making plausible decisions. However, there are still no unified frameworks that incorporate various heuristics, and there has been no comparative evaluation of commonly used heuristics. This paper presents a clear framework to describe various heuristics, and an N‐best algorithm for extracting optimal solutions. The time complexity of this algorithm isO(nNlog2(1 +N)), wherenis the sentence length. The advantage of the N‐best algorithm over the standard beam search algorithm is also discussed. This paper also presents a comparative evaluation of three major heuristics, and proposes a precise and portable rule‐based heuristic. Estimation was done using the aforementioned algorithm and six criteria. The newly proposed heuristic is based upon the Extended Least Bunsetsu (Phrase) Number method.

Read the paper · More papers on PaperTik