Chinese Entity Linking with Two-stage Pre-training Transformer Encoders
Shuai Gong, Xiong Xiong, Shengyang Li, Anqi Liu, Yunfei Liu · 2022
To solve the Chinese entity linking problem, this paper employs a two-stage deep pre-training transformer model. The model is broken down into two stages. The first stage separates the input text and entity information and encodes them in the CN-DBpedia (a popular and large-scale open-source Chinese knowledge base) with bi-encoder. The candidate entities are reranked in the second stage by a deep pre-training transformer encoder (bi-encoder, cross-encoder or polyencoder). To investigate whether the model performs better with different encoders in the re-ranking stage, we examined five Chinese datasets that contained entity link information. Compared to other encoders, the cross-encoder outperforms others, and due to the complexity of Chinese, the difference in performance among different encoders is huge. We noticed that some of the labeled entities in the dataset did not exist in the knowledge base before training, so we added the encyclopedia information to the knowledge base entities. The modified CN-DBpedia contains 9,396,679 entities that could provide better support for the experiment.