Annotation and Classification of Toxicity for Thai Twitter
Sugan Sirihattasak · Institutional Repositories DataBase (IRDB) · 2019
In this study, we present toxicity annotation for a Thai Twitter Corpus as a preliminary exploration for toxicity analysis in the Thai language. We construct a Thai toxic word dictionary and select 3,300 tweets for annotation using the 44 keywords from our dictionary. We obtained 2,027 and 1,273 toxic and nontoxic tweets, respectively; these were labeled by three annotators. The result of corpus analysis indicates that tweets that include toxic words are not always toxic. Further, it is more likely that a tweet is toxic, if it contains toxic words indicating their original meaning. Moreover, disagreements in annotation are primarily because of sarcasm, unclear existing target, and word sense ambiguity. Moreover, we conducted supervised classification using our corpus as a dataset and obtained an accuracy of 0.80, which is comparable with the inter-annotator agreement of this dataset. we also estimate semantic orientation Turney [1] of words to rank words according to toxicity. As the result, we got precision@k for 0.58@40 and 0.41@80. Finally, we launched our demo application for the public feedback and our dataset is available on GitHub.