A Multi-label Classification Algorithm for Non-standard Text

Shihong Chen, Qichao Kuang, Xiaohui Yu, LI Sui-feng, Ruoyao Ding · 2022

To solve the difficulty of extracting semantic features from non-standard text, this paper proposes an algorithm combining deep learning to generate multi-labels and calculating text similarities to copy multi-labels. The algorithm idea is firstly to find the most similar text through calculating the similarities between all texts in the corpus and the target text. If the similarity of the most similar text is higher than the set threshold, the multi-labels of the text will directly copy to the target text; Otherwise, the labels generated by the multi-labels generation algorithm based on ALBERT+CNN are used as the output of the target text. The advantage of our algorithm is that it can not only take advantage of the accuracy and convenience of method to copy labels of similar texts, but also avoid the bias in the similarity calculation of under standardized text. The experimental results show that the algorithm has better performance in dealing with the multi-label text classification problem of non-standard text.

Read the paper · More papers on PaperTik