Enhancing Arabic Text Classification with a Hybrid Word Embedding Method
Eman Aljohani · 2023
In this work, we propose a novel hybrid neural network architecture that combines the contextual sensitivity of Bidirectional Gated Recurrent Units (BiGRUs) with the semantic power of pre-trained FastText word embeddings. By incorporating Term Frequency-Inverse Document Frequency (TF-IDF) weighted features, our method improves feature representation by emphasizing both term significance and document-specific word usage. This combination makes it possible to represent text data in a more nuanced way, particularly for morphologically rich languages. Prior to classification via a dense layer, our model captures both the distinct contextual importance of terms and the deep semantic relationships encoded in the embeddings by concatenating the output of the BiGRU layers with TF-IDF features. The effectiveness of combining different machine learning techniques is highlighted by our results, which show a significant improvement in classification performance over traditional models. The results show that our hybrid approach, which achieved an impressive accuracy result of 0.97, can be a useful tool for difficult language tasks. It provides excellent accuracy as well as insightful information about how algorithmic text analysis is developing.