An Word2vec based on Chinese Medical Knowledge
Jiayi Zhu, Pin Ni, Yuming Li, Junkun Peng, Zhenjin Dai, Gangmin Li, Xuming Bai · 2019
Introducing a large amount of external prior domain knowledge will effectively improve the performance of the word embedded language model in downstream NLP tasks. Based on this assumption, we collect and collate a medical corpus data with about 36M (Million) characters and use the data of CCKS2019 as the test set to carry out multiple classifications and named entity recognition (NER) tasks with the generated word and character vectors. Compared with the results of BERT, our models obtained the ideal performance and efficiency results.