BWSC: BERT-Based N-Gram Enhanced Word Segmentation for Chinese

Wei Hua, Xu Guo · 2024

The word segmentation is a key fundamental task in the field of natural language processing (NLP) and a prerequisite for advanced tasks such as machine translation, sentiment analysis and language recognition. In Chinese language processing, the difficulty of word segmentation is particularly prominent. Because Chinese is based on characters, there are no symbols such as spaces between words to indicate the boundary of words, and “words” are the smallest meaningful language components that can act independently. In practical applications, the semantics of a text are not only dependent on a single word, but also closely related to the larger granularity of the language structure. In the pre-training stage of the model, text is usually divided into small-grained units such as characters or lexical blocks for modeling, and it is easy to ignore the information carried by larger-grained text structures, which may lead to the loss of important semantic information, and the results of word segmentation directly affect the accuracy of subsequent NLP tasks, such as parts-of-speech tagging, named entity recognition, and syntax analysis. In order to improve the accuracy of Chinese word segmentation (CWS), a new model called BERT-based N-gram enhanced word segmentation for Chinese (BWSC) is proposed in this paper. The model combines the contextual coding capabilities of the BERT pre-trained model and is pre-processed by the Jieba word segmentation tool to provide consistent word segmentation criteria. Jieba not only ensures the consistency of word segmentation, but also can automatically identify unregistered words (such as new words and proper nouns), to improve the recognition and processing ability of special words. BWSC enhances the comprehensive expression of character sequences by introducing N-gram features, so that the model can capture the relationship between characters, words and phrases. The Fl score of BWSC in SIGHAN2005 dataset reaches 96.63%. The experimental results show that the combination of Jieba pretreatment and N-gram features significantly improves the CWS effect of the BWSC model, thus providing more accurate input for subsequent NLP tasks. Through the application of BWSC model, the lexical boundary and grammatical structure of text can be captured more accurately, which provides a new technical reference for Chinese NLP research.

Read the paper · More papers on PaperTik