Chinese Government Official Document Named Entity Recognition Based on Albert
Ziqi Xiong, Dezhi Kong, Zhichao Xia, Yankai Xue, Ziyu Song, Peng Wang · 2021
The automated processing of Chinese government documents is in its early stage, and information extraction based on Named Entity Recognition (NER) plays an important role in the automated processing and analysis of Chinese government documents. This paper proposes and implements the pre-trained language model called GovAlbert Based on Albert which the pre-trained language model, which for the processing of Chinese government official documents. We study and analyze NER tasks of the Chinese government official document based on the pre-trained language model, and annotate the Chinese government official documents' Entity recognition corpus, and construct four named entity recognition models based on GovAlbert. The experimental results show that the GovAlbert model for government official document processing has an improved macro-average F1 value (harmonized average of accuracy and recall) than Albert. four named entity recognition models based on GovAlbert in multiple NER tasks of government official documents are all better than the public pre-training model, and through experiments, it has been explored that the GovAlbert-CRF combined model can achieve the best F1 value, so it can be better qualified for the NER tasks of government official documents.