Research on domain text classification method based on BERT
Shaohui Xie, Yangsen Zhang, Zhengyu Hou, Zhenjiang Su, Renjie Wang, Ziyuan He · International Conference on Artificial Intelligence and Intelligent Information Processing (AIIIP 2022) · 2022
Domain texts usually have significant domain features and long text lengths. Text classification models suitable for general fields cannot well meet the task of text classification in specific domains. Therefore, this paper proposes a domain text classification method based on BERT. First, segment the long domain text to obtain the sequence combination, and input it into the BERT pre-training model to obtain the word vector. Then the vector is compressed and encoded to obtain the pooled sequence feature vector. Finally, it is sequentially input to the Encoder layer for domain feature extraction, and the text is divided into various categories in the output layer. The experimental results show that the proposed BERT_VCA model has an average improvement of 1.12% in F1 value compared with the BERT_BASE model in the domain text classification task.