Improving Text Classification by Fusing Linguistic and Semantic Features

Sarang Shaikh, Ehtesham Hashmi, Sule Yildirim Yayilgan, Mohamed Abomhara, Rajendra A. Akerkar · 2025

Text classification remains a fundamental challenge in natural language processing (NLP), with performance often limited by the reliance on either traditional linguistic features or semantic embedding techniques in isolation. This study addresses this limitation by proposing a feature fusion method that integrates traditional linguistic features—such as part-of-speech tags, bag-of-words, TF-IDF, and n-grams—with advanced semantic embedding techniques like word 2 vec and doc 2 vec. The proposed approach aims to capture both syntactic and semantic nuances, enhancing the robustness and accuracy of text classification tasks. To evaluate its effectiveness, the method was applied to five datasets across three critical domains: fake news detection, bloom’s taxonomy classification, and hate speech detection. Key performance metrics, including accuracy, precision, recall, and f1-score, were used to assess the performance of the proposed approach. The experimental results demonstrate that the fusion of linguistic and semantic features consistently outperforms compared to using the each feature type alone, achieving accuracies of $\mathbf{7 9 \%}$ and $\mathbf{6 7 \%}$ for fake news datasets, $38 \%$ and $\mathbf{6 4 \%}$ for bloom’s taxonomy datasets, and $\mathbf{7 0 \%}$ for hate speech datasets. The findings highlight the proposed method’s ability to bridge syntactic and semantic gaps, offering a robust solution for improving text classification performance across diverse NLP applications.

Read the paper · More papers on PaperTik