Design and Implementation of Web Information Extraction Based on HtmlParser
Huang Zhiping · Journal of Jiangxi University of Science and Technology · 2007
The rapid growth of the Web contents increases the need for some automatic tools to help people find the information among the magnanimous information sources such as tides,links,emails,pictures etc.The Web pages expressed by HTML,after analyzed by Internet Explorer,are only suitable for browse,but not for machine process- ing as the way of data exchange.The paper explains how to use HtmlParser to extract hyperlink information from web page,then store in SQL database after cleaning in information detail.