BERT-Guided Pseudo Label Generation for Topical Text Classification
Luepol Pipanmekaporn, Sitthipon Yamvari, Sutwatchai Kamonsantiroj, Kanabadee Srisomboon, Wilaiporn Lee, Akara Prayote · 2025
Topical text classification, which involves organizing text based on its subject matter or topics, faces significant challenges such as limited training data, ambiguous boundaries between closely related topics, and inconsistencies in text content. To address these challenges, this paper introduces a novel approach leveraging BERT, a neural masked language model, to generate pseudo-label for target topics without relying on ground truth data. Our approach uses prompt templates to augment unlabeled text, enabling BERT to predict topic-related words. These extracted words are then used to construct representations for the target topics. To make use of these representations, we introduce two heuristic methods for pseudo-labeling documents. Finally, we train a text classifier from pseudo-labeled documents using a self-training framework to further improve classification accuracy. Experimental results on four benchmark datasets demonstrate that the proposed approach achieves accurate and reliable classification even in the absence of labeled data and outperforms state-of-the-art unsupervised text classification methods across diverse domains.