A Method for Information Extraction from the Web
Gilles Nachouki · 2006
Many data sources are available today on the Web like product catalogs, conference, and multiple directories. The extraction of information from the content is a hard task since they are heterogeneous and dynamic. This paper presents a new method for extracting wrappers and relations from the Web that combines the discovery of similarities in the structures of data that a user wishes to extract from a given Web page, and the generalization of contexts of the extracted data. This method is now implemented in MDSManager our system for data sources Fusion