An Unbalanced Classification Model of Tibetan Medicine Journals with Improved Prompt Learning Method
Yi Gao, Songbin Chen · 2023
In order to solve the problem of low classification accuracy of Tibetan medicine-related journals using pre-training model under the condition of unbalanced sample size. This paper proposes an unbalanced data augmentation for prompt learning based on bert (Unbalanced Data Augmentation for Prompt Learning based on Bert) Based on the BERT model, the unbalanced Tibetan medicine journal data is sent to EDA and SimBERT to generate the balanced data, and then the prompt template is constructed to help understand the downstream tasks, so as to improve the model's extraction of journal subject features. Experiments show that the average subject classification accuracy of the model UDA-PL-BERT is 75%, the accuracy is 97%, the f1 value is 71%, and the recall rate is 60%, which is improved compared with the BiLSTM model, TextCNN model, FastText model and Transformer model. The data modes processed by this method are limited to text form, and the target task is mainly unbalanced multi-classification task of Tibetan medicine-related journal topics, and does not do other dataset classification problems. The UDA-PL-BERT proposed in this paper can effectively improve the extraction and classification performance of unbalanced subject features of Tibetan medicine-related journals.