POS Tagging for Informal Data in Platform X Using Word2Vec and Bidirectional LSTM
Fika Hastarita Rachman, Bain Khusnul Khotimah, Imamah Imamah, M Afifudin Abdullah · 2024
Unstructured and informal text data is often challenging in developing Natural Language Processing (NLP) applications, especially in tasks requiring in-depth linguistic analysis, such as POS Tagging. To address this issue, we propose using a Bidirectional Long-Short-Term Memory (Bi-LSTM) method to perform more accurate POS tagging in Indonesian sentences. This method allows the model to consider the context of the word, thereby improving the model's understanding of the meaning and function of the word in the sentence. Text preprocessing includes data cleaning, case folding, tokenization, and normalization, which are used before the feature extraction process. Word embedding with Skip-Gram is a feature extraction method that converts words into vectors and captures semantic meaning in the Indonesian context. The Dataset is divided into training and test, combined with 5-fold cross-validation to ensure more stable and accurate results. The results showed that this approach produced an FI score of 94.52% in the best model, indicating a high level of accuracy in word class recognition. With these results, the Bi-LSTM method combined with Skip-Gram word embedding shows excellent potential for overcoming the challenges of linguistic analysis in Indonesian informal texts, which can be applied to other NLP tasks.