Chinese Real-word Error Automatic Detection and Correction Based on Confusion Set and Generalization Model
Haitao Wang, Xinyu Cao, Liangliang Liu, Dezhi Gu · 2020
With the development of informatization, there are many errors in Chinese texts, and one kind of error is called “Real-word Error” which like English Real-word Error. The wrong word itself is also a right word in dictionary which confirm by the specific context. While using data resource of corpus to proofread text, data Sparseness will bring the miscarriage of justice problem. In view of the above problem, this paper proposes a Chinese real word error detection and correction method based on confusion set and generalization model. The method constructs the Chinese word confusion sets and trigram, and combines N-gram model and Bayesian model to preliminarily screen the current word and further validate on the probability of the current word appearing in the context respectively. In addition, the method alleviates the problem of false positive caused by sparse corpus data by synonym generalizing the context features. The proposed method combines automatic error-detecting and automatic error-correction together. Experimental results show that the proposed method can effectively find Real-word Error in Chinese text, and has higher recall rate, detecting accuracy rate and correcting accuracy rate and can give the correction lists.