Improve the Performance of the Webpage Content Extraction Using Webpage Segmentation Algorithm
Fu Lei, Yao Meng, Yu Hao · 2009
In this paper, we present a method using Webpage segmentation algorithm to improve the performance of the Webpage content extraction. The traditional methods often depend on parsing the DOM tree of the Webpage and judging each node of the DOM tree to determine which node is the text node, this kind of method has a potential problem, it sometimes throws part of the content away because of its local judgement strategy. But our method which is based on the VIPS (vision-based page segmentation) algorithm, can solve the problem satisfactorily, it can extract the content according to the coordinate information of the block and help the traditional method to recall the lost part of the content.