The Effect of Feature Extraction Based on the Frequency of Emotionally Loaded Words using Part-of-Speech Tagging in Sentiment Analysis

Alyza Rahima Pramduya, Shafa Amira Qonitatin, Budi Juarto · 2025

Sentiment analysis plays a vital role in understanding public opinions. In this paper, we propose a methodology to improve sentiment classification models with features extracted using Part-of-Speech (POS) tagging. The feature extraction process involves building a vocabulary of frequently occurring words associated with specific parts of speech in both positive and negative sentiment datasets. These vocabularies are then used to calculate the frequency of occurrences in both training and testing datasets. We applied three traditional machine learning algorithms: Decision Tree, Random Forest, and K-Nearest Neighbors (KNN) to evaluate the impact of these features. The dataset used in this study was collected from YouTube comments, focusing on public sentiment in Indonesia regarding the relocation of the capital city to Ibu Kota Nusantara (IKN). F1-score was used as the performance metric for evaluation. The experimental results demonstrated that models incorporating POS-tagged features showed improved performance compared to models without POS tagging. This suggests that POS tagging contributes to enhancing sentiment classification.

Read the paper · More papers on PaperTik