Burmese Word Segmentation Method and Implementation Based on CRF
Chang'e Ma, Jian Yang · 2018
Burmese belongs to the languages whose writing system has no delimiters to mark word boundaries. However, related works on Burmese word segmentation are still at the initial stage. This paper aims to fill the blank by employing CRF model to the task. The performance of the CRF method is evaluated with confidence, precision of segmentation. We prepared an experimental database of 5,000 sentences, which were manually segmented by Burmese experts. After the 6-fold cross-validation of the experimental data set, the experimental results show that the average confidence level of the CRF method is 93.4%, which is greater than the threshold, and the average value of the F1 is 93.0%. Therefore, the CRF segmentation method satisfies the requirements for developing a Burmese speech synthesis system.