Study of Automatic Knowledge Extraction in Specific Chinese Language Domain
Suxiang Zhang, Lei Li, Yixin Zhong · 2006
The paper presents hierarchy bootstrapping as an alternative approach to learning from a large quantity of unlabeled data in the Chinese language domain. It advocates using a small amount of seed information and a large collection of easily-obtained unlabeled data. Hierarchy bootstrapping initializes a learner with seed information; then it iterates applying the learner to calculate for the unlabeled data. Two case studies of this approach are presented in order to solve the problem of automatic knowledge extraction in information extraction (IE) systems. The first algorithm makes use of seed words and seed patterns to build a learner, which extracts more characteristic words using scalar clusters method. These characteristic words have semantic similarity with seed words. Then more extraction patterns could be learned automatically and added to the knowledge database by using the second algorithm, they are a foundation for analysis of IE. Experimental results are promising