A Method Combining Text Classification and Keyword Recognition to Improve Long Text Information Mining

Lang Liu, Yuejiao Wu, Lujun Yin, Junxiang Ren, Ruiling Song, Guoqiang Xu · 2022

Using a pre-trained language model for long text classification will be limited by the input length of the pre-trained language model. At the same time, we cannot effectively utilize all the text information in the long text. While using TextCNN for long text classification will be limited by the input word embedding based on a specific task, and it is unable to fully understand the semantics of the current task. In order to better understand the semantic information of the current task, we will use TextCNN with special word embedding for long text classification, we use the keyword data of the current task to finetune the pre-training task to obtain better word embedding representations. On the one hand, using the current word embedding representation as inputs to the classification task better supports the classification task, and on the other hand, it further finetunes the pre-training task to obtain a better word embedding representation. At the same time, the keyword data can be better expanded. Compared with the direct use of TextCNN for long text classification, long text classification combined with keywords is more efficiency and achieves the effect close to BERT. Experiments are carried out on the three datasets, ChnSentiCorp, NLPCC14-SC and business opportunity recommendation(BOR). The model accuracy is relatively high, with the increase of 1.8%, 1.03% and 4.03%, the relative increase of F1-score by 1.8%, 1.03% and 5.04% respectively. Experiments show that, without changing the efficiency of the TextCNN text classification model, we keep the integrity of the input text, the pre-training task finetuned by keyword discovery task enables the word embedding to have better understanding of current semantic expression. The text classification model combined with keywords discovery task by co-training can effectively improve the classification effect on different datasets.

Read the paper · More papers on PaperTik