Named Entity Recognition of Chinese Legal Text Based on BERT
Huawei Lu, Yinan Peng · 2022
With the deepening of national judicial reform, how to combine artificial intelligence with judicial work has become a research emphasis in judicial intelligence research. Aiming at polysemy problem in Chinese, and characteristics of Chinese legal text that complicated context, professional, and diverse types of entities, we design a named entity recognition method for Chinese legal text based on BERT. Firstly, we build a Chinese legal text corpus, and utilize the corpus to domain pre-train the pre-training model BERT to make it perform better for named entity recognition task in Chinese legal text. Secondly, we design the method of adding clue words and the method of replacing synonyms for training data augmentation to increase the diversity of training data. Finally, we combine the BERT model after domain pre-training with a CRF layer, and utilize the datasets after data augmentation to train it. Meanwhile, we utilize adversarial training in the training process to improve the generalization ability of the model. To verify the performance of our model, we conduct experiments on the competition data set of the information extraction track of the China Legal Intelligence Technology Evaluation Competition, which proves the feasibility and effectiveness of the method.