Boosting Text Classification Performance for Unlabeled Data with Semi-Supervised Learning
Tanvina Khondokar, Saiful Islam, Tasmia Ishrat Alam Chadni, Iftakhar Ali Khandokar · 2023
Text classification has always been an engrossing problem in the field of Natural Language Processing. To build a model on a specific domain, it needs to gain the ability to distinguish text that is related to that domain and others that are not. To make the model achieve this type of domain knowledge, we need to provide a large amount of labeled data within the domain. Having said that, labeling such an amount of data is challenging, as it demands hard labor, a tremendous amount of time, and cost. These could slow the progress of any language or domain model, achieving the performance that is expected. In this paper, we have attempted to address the above-discussed problem, proposing a semi-supervised performance-boosting technique for large-volume unlabeled datasets. The target of this work is to demonstrate the semi-supervised un-labeled data labeling mechanism that can boost the accuracy of text classification models. We have developed an ensemble semi-supervised labeling technique that utilizes the weighted output of different feature-based predictive classifiers to classify unlabelled datasets which makes the labeling process more reliable than unsupervised approaches. The proposed approach overcomes the above-mentioned challenges. After evaluating our approach, we observed that our Semi-Supervised boosting technique enhanced the performance of the predictive models to 20% approx.