A Preliminary Study on Fundamental Thai NLP Tasks for User-generated Web Content
Anuruth Lertpiya, Teerapat Chaiwachirasak, Nattasit Maharattanamalai, Theerapat Lapjaturapit, Tawunrat Chalothorn, Nutcha Tirasaroj, Ekapol Chuangsuwanich · 2018
Existing literature on Thai NLP often focuses on formally written texts with near-perfect spellings and boundaries between words or sentences. Such assumptions, however, do not hold in real-world NLP tasks, especially when dealing with User-generated web content (UGWC). So far, existing NLP research works on actual web data have been limited, making it unclear whether and how existing techniques can be applicable to UGWC. In this paper, several basic Thai NLP algorithms (word segmentation, sentence segmentation, word error detection, word variant detection, name entity recognition) are re-investigated and benchmarked against real-world, practical UGWC data set. The difference in performance between our data set and others are compared as a guidance for future research. Our baseline sentence segmentation on UGWC data set yields an average F-measure of 0.77. For name entity recognition and word variant / error detection tasks, our system yields the accuracy of 0.93 and 0.53, respectively.