R-Extractor: A Method for Data Extraction from Template-Based Entity-Pages

Edimar Mânica, Carina F. Dorneles, Renata Galante · 2017

The challenges in Big Data start during the data acquisition, where it is necessary to transform non-structured data into a structured format. The Big Data Era brought a new challenge to data acquisition: the need for eliminating user intervention. In this context, the domain-centric data extraction (DCDE) methods arose through replacing user intervention with content redundancy. The DCDE methods extract the attribute values of entities of the web that are restricted to a specific application domain. A research gap was identified from various analyzed methods, i.e., the DCDE methods are not effective in extracting attribute values that are differently presented within the same website. We have responded to this research gap by proposing the R-Extractor method, which extends a state-of-the-art DCDE method by adding a reinforcement stage for the treatment of attribute values that are differently presented within the same website. The R-Extractor method re-analyzes the extraction rules to identify those that can be combined to extract the values of a given attribute from all the pages that describe entities on a website. This identification is based on a novel score function that takes into account different features of the extraction rules. We carried out experiments on a dataset with more than 50k web pages from different real-world websites of a wide range of application domains. The R-Extractor method reached 98% of precision. Our method was compared with two baselines (an XPath-based method and a tree-based method) and outperformed them with an increase in precision up to 14%.

Read the paper · More papers on PaperTik