Hybrid Feature Fusion Learning Towards Chinese Chemical Literature Word Segmentation
Xiang Li, Kewen Zhang, Quanyin Zhu, Yuanyuan Wang, Jialin Ma · IEEE Access · 2021
The rapid increase in the number of chemical science literature has brought challenges to researchers in search and data analysis. For many chemical scientific literature, extracting information from text and using knowledge is the focus of research. However, the existing Chinese text word segmentation methods have low recognition rate for chemical terms. The reason is that the addition of many new vocabulary and mixed professional vocabulary of Chinese and English brings challenges to word segmentation. In this paper, we propose a word segmentation method of Chinese chemical literature based on hybrid feature fusion learning Model(HFFLM). HFFLM first establishes chemical science corpus (Chem-pku) to train Chinese word segmentation (CWS) tasks. In addition, HFFLM uses BiLSTM and CNN to extract document features and fuse them. Then, HFFLM combines boundary features to construct conditional random field to train the end-to-end CWS model. In the end, HFFLM makes visual analysis of the word segmentation results. The experimental results indicate that HFFLM has high accuracy and recall rate, and is suitable for chemical industry vocabulary extraction with mixed Chinese and English.