Chinese Characters and Pinyin: A Model with Two Parallel Feature Extractors for Chinese Entity Recognition

Yuhui Zhai, Shengwei Tian, Long Yu, Bo Wang, Tiejun Zhou, Yu Wang, Mandeng Gao · Research Square · 2022

Abstract The purpose of Named Entity Recognition (NER) is to identify and mark entities with specific meanings in a text. Compared with English NER, Chinese NER is blurry about the boundaries of entity classes because there is no clear separator between Chinese characters and Chinese entities are composed of several characters with different lengths. For Chinese NER, traditional methods only focus on Chinese characters, ignoring the important role of pronunciation. But for these models considering pronunciation, they put pronunciation and characters together for feature extraction. In this paper, we propose a Model with Two Parallel Feature Extractors. It uses a new Pinyin embedding layer that can handle characters except Pinyin and it uses Pinyin Encoder and Word Encoder to obtain the features of Pinyin and characters respectively and then features are fused through TextCNN. Compared with the previous model, this model is not as big as BERT, and it can get good results without additional data type training. We used four datasets: Resume, CCKS2019, CLUENER2020 and MSRA to test our model and it showed a good result, which proved the validity of our model.

Read the paper · More papers on PaperTik