Patent Text Classification Study Based on Bi-LSTM-A Model

Jinhong Wu, Caiyun Huang, Yongyue Chen · 2020

At present, the classification of patent text is still processed manually, but in the classification of patent IPC level, the similarity of text is getting higher and higher, and the difficulty of subjective classification and categorization is increasing one layer at a time, which is time-consuming and laborious. [Method] According to the characteristics of the patent hierarchy with different text similarity and difficult to distinguish, the model improves the classification effect by encoding the input text and introducing an attention mechanism to assign different weight values to different text. Finally, the validity of the model is verified by experimenting with the corresponding data set from the smart bud and comparing it with the traditional classification methods: decision tree, SVM. [Results] The classification accuracy rate of all levels in the Bi-LSTM-A model is good, and most of the results are above 94%, which is effective for automatic classification of patent texts, and the best performance is obtained in experiments with high text similarity. [Limitations] In this paper, only part of the data of the section, class and group were collected for the experiment, and not enough data were available for the experiment. If the classification effect is judged simply on the basis of the correct rate, the traditional classification algorithm can also perform automatic classification better. However, considering the data volume, data balance and different text similarity, the Bi-LSTM-A model proposed in this paper has a higher classification accuracy under the data imbalance and high text similarity, and the robust model is better.

Read the paper · More papers on PaperTik