A Method for Building Chinese Domain Lexicon Based on New Words Recommendation
Jinjian Duan, Meiqing Wang, Yibo Guan, Qiu Lin · 2022
Professional domain lexicon construction is the basis for Domain-specific Natural Language Processing. Current new words detection algorithms struggle with migration and sparsity and are unable to facilitate the building of domain lexicons. In this paper, a method for building a domain-specific lexicon based on new words recommendation is proposed. A model of the domain lexicon construction process was developed, including domain corpus pre-processing, new words detection, saliency evaluation, and new words recommendation. The problem of the sparsity of unsupervised new words discovery algorithms is solved by using Enhanced Mutual Information and Branch Entropy to identify small word frequency words and NC-value to identify domain words. According to experiments on a guided domain dataset, the method is effective at enhancing the precision of domain new words recommendations, lowering the complexity of expert work, and has significant implications for the building of a domain-specific lexicon.