Automatic Detection and Correction for Chinese Misspelled Words Using Phonological and Orthographic Similarities

Tao-Hsing Chang, Hsueh‐Chih Chen, Yuen‐Hsien Tseng, Jian-Liang Zheng · 2013

How to detect and correct misspelled words in documents is a very important issue for Mandarin and Japanese. This paper uses pho-nological similarity and orthographic similar-ity co-occurrence to train linear regression model. Using ACL-SIGHAN 2013 Bake-off Dataset, experimental results indicate that the detection F-score, error location F-score of our proposed method for Subtask 1 is 0.70 and 0.43 respectively, and the correction ac-curacy of the proposed method for Subtask 1 is 0.39. 1

Read the paper · More papers on PaperTik