Morphological Analysis for Unsegmented Languages using Recurrent Neural Network Language Model
Hajime Morita, Daisuke Kawahara, Sadao Kurohashi · 2015
We present a new morphological analy-sis model that considers semantic plausi-bility of word sequences by using a re-current neural network language model (RNNLM). In unsegmented languages, since language models are learned from automatically segmented texts and in-evitably contain errors, it is not apparent that conventional language models con-tribute to morphological analysis. To solve this problem, we do not use language mod-els based on raw word sequences but use a semantically generalized language model, RNNLM, in morphological analysis. In our experiments on two Japanese corpora, our proposed model significantly outper-formed baseline models. This result indi-cates the effectiveness of RNNLM in mor-phological analysis. 1