Research on Multi-Label Classification of Chinese Text Based on Word-Label Probability and DistilBERT

Hong Ping Zhao, Weiquan Tian, Ding An, Shan Mi · 2024

Multi-label text classification is a fundamental task in the field of natural language processing. Currently, there are issues in the Chinese multi-label text classification tasks, such as insufficient extraction of text label features and a lack of learning about the correlations between labels. To address these challenges, we proposed a Multi-label Text Classification Model based on DistilBERT and Word-Label Probabilities (DBWL). The model consists of a metric space module that integrates word-label probabilities and a deep neural network module based on DistilBERT. Firstly, the trained Labeled Latent Dirichlet Allocation (Labeled-LDA) model is utilized to obtain the probabilities of words associated with labels, which are then used to calculate the text mapping vectors and map them into the same metric space as the labels. Similarity measurement is employed to extract features between text-label pairs and label-label pairs. Then, in the deep neural network module, DistilBERT is employed to jointly embed texts and labels, capturing rich semantic information about text labels, followed by feature extraction using TextCNN. Finally, the features extracted from both modules are weighted and fused for label prediction. Experimental results on two benchmark datasets demonstrate that the proposed model outperforms current mainstream multi-label classification models in terms of key performance metrics.

Read the paper · More papers on PaperTik