Histogram Equalization on Statistical Approaches for Chinese Unknown Word Extraction
Bor-Shen Lin, Yi-Cong Chen · 2011
With the evolution of human lives and the spread of information, new things emerge quickly and new terms are created every day. Therefore, it is important for natural language processing systems to extract new words in progression with time. Due to the broad areas of applications, however, there might exist the mismatch of statistical characteristics between the training domain and the testing domain, which inevitably degrades the performance of word extraction. This paper proposes a scheme of word extraction in which histogram equalization for feature normalization is used. Through this scheme, the mismatch of the feature distributions due to different corpus sizes or changes of domain can be compensated for appropriately such that unknown word extraction becomes more reliable and applicable to novice domains. The scheme was initially evaluated on the corpora announced in SIGHAN2. 68.43 % and 71.40 % F-measures for word identification, which correspond to 66.72%/32.94 % and 75.99%/58.39 % recall rates for IV/OOV, respectively, were