Research on Chinese word segmentation model based on RoBERTa and conditional radom field

Wenfeng Cao · 2023

This paper proposes a Chinese word segmentation model based on RoBERTa and Conditional Random Field. The Chinese word segmentation model based on RoBERTa and CRF deals with Chinese word segmentation in the way of sequence labeling. The training of the model takes the character sequence of the sentence as the input, which is converted into number according to vocabulary and sent into RoBERTa layer. The output results of RoBERTa layer were converted into the probability distribution of word position label by linear transformation and softmax activation function. The predicted word position label sequence was obtained by Viterbi decoding operation. After experiments, the accuracy rate, precision rate, recall rate and f1 scores of RoBERTa + CRF model on SIGHAN's PKU and MSR corpora are 96.33%, 95.75%, 95.00% and 95.37% respectively (pick the highest score), which has preliminary application value and certain room for improvement.

Read the paper · More papers on PaperTik