Sentiment classification model for Uyghur language texts based on BERT BiLSTM

Feierdun Keremu, Liu Li · 2024

With the rapid expansion of the Internet, the proliferation of online text presents significant challenges to natural language processing (NLP) methodologies. In recent years, deep learning has emerged as a focal point within the realm of NLP. Nevertheless, sentiment analysis of Uyghur text has encountered mounting hurdles due to the scarcity of Uyghur language corpora and the intricate nature of word formation and morphological affixes inherent in the language. To address the complexities of extracting nuanced emotional information from Uyghur text, this paper proposes a novel text sentiment classification model that combines the advanced Bidirectional Encoder Representations from Transformers (BERT) with Bidirectional Long Short-Term Memory (Bi-LSTM). This pioneering model marks the first attempt at applying such techniques to Uyghur text sentiment classification. Leveraging the BERT network, this model constructs word embeddings for the Uyghur text, and then utilizes Bi-LSTM to capture both contextual and localized key features. Subsequently, the sentiment category of the text is determined using a softmax classifier. In this study, Uyghur text data was sourced from various online platforms, including Uyghur news websites and forums, to compile a comprehensive Uyghur language corpus. Preprocessing techniques such as deduplication and denoising were applied to refine the corpus. The proposed BERT-BiLSTM model and algorithm underwent rigorous evaluation for sentiment classification performance using the constructed Uyghur text dataset. Results indicate that, in comparison to alternative methods for text sentiment classification, the proposed model achieves superior accuracy, recall, and the comprehensive evaluation metric F1 score.

Read the paper · More papers on PaperTik