LexToPlus: A Thai Lexeme Tokenization and Normalization Tool
Choochart Haruechaiyasak, Alisa Kongthon · 2013
The increasing popularity of social media has a large impact on the evolution of language usage. The evolution includes the transformation of some existing terms to enhance the expression of the writer’s emotion and feeling. Text processing tasks on social media texts have become much more challenging. In this paper, we propose LexToPlus, a Thai lexeme tokenizer with term normalization process. Lex-ToPlus is designed to handle the intentional errors caused by the repeated characters at the end of words. LexToPlus is a dictionary-based parser which detects existing terms in a dictionary. Unknown tokens with repeated characters are merged and removed. We performed statistical analysis and evaluated the performance of the proposed approach by using a Twitter corpus. The experimental results show that the proposed algorithm yields an accuracy of 96.3 % on a test data set. The errors are mostly caused by the out-ofvocabulary problem which can be solved by adding newly found terms into the dictionary. 1