Burmese word segmentation with Character Clustering and CRFs

Myat Lay Phyu, Kiyota Hashimoto · 2017

Word segmentation is one of the most fundamental processes for most natural language processing tasks. In particular, languages with no word boundary in writing such as Chinese, Japanese, Korean, Thai, and Burmese need it. However, the Burmese language still waits for a technique with good performance. In this paper, we propose a new technique for Burmese word segmentation employing the idea of Character Clustering for Conditional Random Fields. Character clusters are groups of some inseparable characters due to language characteristics. We proposed a set of 29 types of Burmese Character Clusters (BCCs) as rules, and Conditional Random Fields is applied as a sequential labelling machine learning method. We compared our proposed method with CRF without BCC and Syllable-based CRFs. The result shows that our proposed method achieved the highest performance.

Read the paper · More papers on PaperTik