Extraction of Informative Blocks from Web Pages Based on VIPS
Cunhe Li, Dong Juan, Juntang Chen · 2010
Apart from the main content blocks, almost all web pages on the Internet contain such blocks as navigation, copyright information, privacy notices, and advertisements, which are not related to the topic of the web page. We call these blocks noisy blocks, and call the main content blocks informative blocks. The information contained in the noisy blocks can seriously harm Web mining and searching. So discriminating informative blocks from the noisy blocks and then extracting the information contained in the informative blocks is an important task. In this paper, we propose a method that utilizes both the visual features and semantic information to extract topic information. We first partition a web page into semantic blocks using improved vision-based page segmentation. The visual and the semantic information are extracted to form the feature-vector of the block. Secondly the blocks with similar content structures and spatial structures are clustered by means of similarity computation. After clustering blocks with similar structures, our algorithm identifies informative clusters. The blocks contained in the cluster are informative blocks. Finally we show that a kNN classifier, taking into account the proposed method, clearly outperforms the same classifier using extract informative block arithmetic.