Algorithm of Part-of-Speech Tagging of Corpus based on HMM Model
Hui Li, Zhijing Wu · 2024
Part-of-speech tagging is a basic work in the field of information processing, and the research results can be directly integrated into many practical applications such as information extraction, information retrieval and machine translation. When using statistical-based HMM for part-of-speech tagging, the state set is the part-of-speech tagging set, and the output value (observation value symbol set) is the word in part-of-speech tagging. Based on HMM model, this paper focuses on the parts-of-speech tagging algorithm of corpus, and carries out experimental research and results analysis. Corpus part-of-speech tagging algorithms mainly include smoothing algorithm and Viterbi algorithm. Smoothing algorithm further adjusts the probability distribution of parameter model to ensure that each probability parameter is not zero. Viterbi algorithm is used to obtain the best part of speech sequence of word sequences, and to solve the tagging problem of double class words. The experimental results show that the closed experiment is more efficient than the open experiment, and the HMM model is better than the MEM model in both the closed experiment and the open experiment.