A Machine Learning Approach for the Classification of Methamphetamine Dealers on Twitter in Thailand

Punnavich Khowrurk, Rachada Kongkachandra · 2020

This research presents a method to classify messages from Twitter (tweet) related to Methamphetamine. The messages are classified into three classes: normal, seller, buyer. The models presented in this research are Multinomial Naive Bayes, Multi-Class LSTM, and Hierarchical LSTM. Model training uses a balanced and imbalanced dataset. The text used for Model training is tokenized from four tokenizers: Tlex+, Lexto+, Attacut, and Deepcut. To study the model performance's effect, we divide the data with a different dataset and tokenizer. The results showed that all models could classify the messages into the three classes. The most effective model built from a balanced dataset is the Hierarchical LSTM model using the Lexto+ Tokenizer provides the highest Accuracy, and the most effective model build from an imbalanced dataset is the Multi-Class LSTM model using the Lexto+ Tokenizer. This model gave the highest Accuracy, but the Fl-Score of the Hierarchical LSTM model gave better Accuracy in each class.The creation of a text classification model related to Methamphetamine uses Twitter messages. Most of them are Thai grammatical errors and has many slang usage. We found that Lexto+ is the best tokenizer to build a model. However, it is not much different from other tokenizers. On the other hand, the best dataset to build the model is a balanced dataset that significantly affects model performance.

Read the paper · More papers on PaperTik