Instance-based wrapper induction
Raymond Kosala, Hendrik Blockeel · Lirias · 2000
Information extraction (IE) is the problem of transforming a collection of documents into information that is more readily digested and analyzed. With the popularity of the Web and the tremendous increase in the amount of data available, there is an increasing need for (semi-) automatic IE systems that are able to extract information from (semi-) structured documents. Building IE systems manually is not feasible and scalable for such a dynamic and diverse medium as the Web. Another problem is the difficulty in porting IE systems to new applications and domains if done manually. To address these difficulties there has been increasing interest in applying machine learning (ML) techniques for Information Extraction. In this paper, we describe...