Improving Thai word segmentation with Named Entity Recognition

Sayan Tepdang, Choochart Haruechaiyasak, Rachada Kongkachandra · 2010

Segmenting words in Thai language is a very difficult task since there is no distinguished clue such as blank, period and other punctuations as in English. Several previous researches employed dictionary as the main resource for consideration. However there still exist two problems including ambiguous words and unknown words. These unknown words can be categorized into two groups, -i.e., newly defined words and named entities. This paper presents an approach for improving the performance of Thai word segmentation by merging Named Entity Recognition (NER) to the Thai word segmentation. The Conditional Random Fields (CRFs) algorithm is applied for training and recognizing Thai named entities. The prefixes and suffixes of Thai named entities are selected as main features for learning the models. The performance evaluations are experimented by using the Thai standard word segmentation corpus, namely BEST2009, which consists of 5 million words. Various word-level grams (i.e., three, five and seven) are also employed to construct the Thai NER models. The experimental results show that the 7-gram NER model provides the best performance. Merging the proposed NER model to the Thai word segmentation called TLex (Thai Lexeme Analyzer) can improve the performance measured by F1-measure from 92.39% to 93.96%.

Read the paper · More papers on PaperTik