A BiLSTM-Based Sentiment Analysis Scheme for Khmer NLP in Time-Series Data
Sokleng Prom, Panharith Sun, Neil Ian Cadungog-Uy, Sa Math, Tharoeun Thap · 2024
This research studies the extraction of sentiments from the Khmer language through machine learning and deep learning methods, specifically the Bidirectional Long Short-Term Memory (BiLSTM) network. To achieve this, it employs a quantitative approach to analyze the result from the different training classifiers, including Support Vector Machine (SVM), Naïve Bayes (NB), Random Forest (RF), K-Nearest Neighbors (K-NN), and the BiLSTM neural network. The study also utilized a pre-trained BERT model on the Khmer language as the embedding model combined with applying preprocessing techniques, such as data cleaning and word segmentation before the classifications. The performance of the models is evaluated using accuracy, precision, recall, and F1-score. The findings reveal that BiLSTM with the contextual embeddings from BERT, achieves the greatest performance resulting in an accuracy rate of 86%, outperforming the traditional machine learning algorithms in classifying Khmer text sentiments into negative, neutral, and positive classes.