Weakly Supervised Chinese Short Text Classification Algorithm Based on ConWea Model
Shoujin Wang, Chengyu Duan, Yuanjiao Yang · 2022
The version sort algorithm based on full supervision needs to use a large amount of label data, and the labeling task of text data is time-consuming and difficult to label. Finally, use high-quality pseudo-label data to train a Chinese brief version systematics model. Experiment on the THUCNews news headlines dataset. The test outcome illustrate the performance of this algorithm is slightly stronger than the mainstream semi-supervised classification algorithm when only a small amount of label data is used, and it is not inferior to the general fully supervised classification algorithm. which provides a better solution for unlabeled data classification tasks.