Part-of-Speech Tagging of Corpus based on BiLSTM-CRF Model

Jinxia Yue · 2025

POS tagging is the basis of natural language processing tasks. Compared with traditional machine learning methods, deep learning models can automatically extract features from text and learn deep implicit information, which not only eliminates the tedious manual feature extraction process, but also improves the accuracy of features, showing excellent performance on large-scale corpora. Aiming at the problem that BiLSTM cannot predict the relationship between labels, this paper introduces the CRF framework and focuses on the BILSTM-CRF algorithm. BiLSTM can learn the context information of target words at the same time, while the conditional random field layer can obtain the label information of the sentence layer through training and learning. The experimental environment and training parameters were designed, and the baseline model associated with the BiLSTM-CRF was selected for comparison experiment. It was concluded that the BiLSTM-CRF model had obvious advantages over the baseline model. The research results not only improve the performance of POS tagging, but also expand the application field of deep learning model.

Read the paper · More papers on PaperTik