Automatic Chinese Confusion Words Extraction Using Conditional Random Fields and the Web
Chun‐Hung Wang, Jason S. Chang, Jian-Cheng Wu · 2013
A ready set of commonly confused words plays an important role in spelling error detec-tion and correction in texts. In this paper, we present a system named ACE (Automatic Con-fusion words Extraction), which takes a Chi-nese word as input (e.g., “不脛而走”) and au-tomatically outputs its easily confused words (e.g., “不徑而走”, “不逕而走”). The purpose of ACE is similar to web-based set expan-sion – the problem of finding all instances (e.g. “Halloween”, “Thanksgiving Day”, “Inde-pendence Day”, etc.) of a set given a small number of class names (e.g. “holidays”). Un-like set expansion, our system is used to pro-duce commonly confused words of a given Chinese word. In brief, we use some hand-coded patterns to find a set of sentence frag-ments from search engine, and then assign an array of tags to each character in each sentence fragment. Finally, these tagged fragments are served as inputs to a pre-learned conditional random fields (CRFs) model. We present ex-periment results on 3,211 test cases, showing that our system can achieve 95.2 % precision rate while maintaining 91.2 % recall rate. 1