Research on Multi-Label Text Classification Based on Multi-Channel CNN and BiLSTM

Shoujin Wang, Yuanjiao Yang, Xin Meng · 2022

Natural language processing research is moving in a significant improvement toward multi-label text classification. Multi-label text classification is presently used intensively in practical disciplines. In the text classification task, due to the professional and complex diversity of the text, it is difficult to fully represent the semantic of the text only by relying on the existing word vector representation method, which leads to the low accuracy of the classification task. In order to even get dynamic semantic information of text, this dissertation proposes a hybrid neural network model that incorporates use of Bert's pre-trained language model. Bert model has achieved remarkable results in text classification task. In order to analyze the text content, multi-channel CNN and BiLSTM are proposed for feature fusion. Simple CNN (Convolutional Neural Network) and RNN (Recurrent Neural Networks) cause problems such as loss of key feature information and poor classification performance when processing classification tasks. Multi-channel CNN uses different convolution kernels to obtain local semantic features of text, and BiLSTM model uses gating mechanism to obtain global information of text context. The ultimate vector representation of the text proposed in this study is the result of superimposing the two components mentioned above, and Sigmoid is exploited for classification. The suggested multi-label text classification method provides a stronger classification effect than existing models, thus according experimental results. It can be found that our model has achieved the maximum value in three indexes, with Precisionmicro, Recallmicroand F1microreaching 0.9455, 0.9181 and 0.9315.

Read the paper · More papers on PaperTik