Vision-based Web Data Records Extraction
Wei Liu, Xiaofeng Meng, Weiyi Meng · 2006
This paper studies the problem of extracting data records on the response pages returned from web databases or search engines. Existing solutions to this problem are based pri-marily on analyzing the HTML DOM trees and tags of the response pages. While these solutions can achieve good re-sults, they are too heavily dependent on the specifics of HTML and they may have to be changed should the re-sponse pages are written in a totally different markup lan-guage. In this paper, we propose a novel and language in-dependent technique to solve the data extraction problem. Our proposed solution performs the extraction using only the visual information of the response pages when they are rendered on web browsers. We analyze several types of vi-sual features in this paper. We also propose a new mea-sure revision to evaluate the extraction performance. This measure reflects perfect extraction ratio among all response pages. Our experimental results indicate that this vision-based approach can achieve very high extraction accuracy.