A Semantic Annotation Tool to Extract Instances from Korean Web Documents.
Hai-Tao Zheng, Bo‐Yeong Kang, Sang-Ok Koo, Hee-Chul Choi, Kwang-Sub Kim, Hong‐Gee Kim · 2006
Although there has been extensive research on developing semantic annotation tools recently, only few systems support automatic information extraction. In this paper, we propose a semantic annotation system named SARM, which has an automatic instance extraction module based on two machine learning techniques, Bayesian Classifier and Support Vector Machine. SARM has been tested to make a Korean Restaurant ontology evolve by automatically extracting instances from Web documents in Korean. The automatic instance extraction module can accelerate the annotation work which is very time-consuming and involves a lot of human labor. We describe the implementation of our system and also compare the performances of the two machine learning methods we used.