Text Mining From Site Invariant and Dependent Features For Information Extraction Knowledge Adaptation
Tak-Lam Wong, Wai Pang Lam · 2004
We develop a framework which can adapt previously learned information extraction knowledge from the source Web site to new unseen sites. Our framework also makes use of data items previously extracted or collected. Site invariant features are derived from the previously learned extraction knowledge and previously collected items. Multiple text mining methods are employed to automatically discover machine labeled training examples for the new site. Both site invariant and site dependent features of these machine labeled training examples are used to learn the new extraction knowledge. Extensive experiments on real-world Web sites have been conducted to demonstrate the effectiveness of our framework.