Research on Computerized Automatic Word Segmentation of Chinese Stylistic Words

Menghan Xu · 2023

Natural language processing involves various aspects such as phonology, lexicon, syntax and semantics. In the case of Chinese, words are the smallest linguistic units that can be used independently, and they are also the basic units of Chinese natural language processing. In order to promote the computer’s mining and utilization of the relevant data in the form of natural language text in different stylistic, and improve the recognition and classification ability of the word segmentation system for specific function or domain words, this paper investigates the Chinese stylistic words based on computerized automatic word segmentation. According to the sampling method, a certain number of language materials are selected from different types of modern Chinese, and five word segmentation models of NLPIR- ICTCLAS, Jieba, LTP, THULAC and PKUSeg are used to divide these texts, and then the segmentation model which best fits the characteristics of stylistic words is selected. Then, based on the word segmentation results of the model and the theories related to linguistics, a certain number of stylistic words are extracted. This research can promote the machine’s mastery of natural language rules, and play an important role in the classification of language materials, recognition and retrieval of subject information, and contextual semantic analysis. And besides, it can improve the efficiency and accuracy of computer’s recognition of unregistered words of specific text and the accuracy of machine translation, which are of great value in both Chinese natural language understanding and natural language generation.

Read the paper · More papers on PaperTik