Design and Implementation of Web Information Extraction System Based on HTMLParser
Xinghua Hou · Journal of Liaoning University of Petroleum & Chemical Technology · 2006
The rapid growth of the Web contents increases the need for some automatic tools to help to find the exact information among the magnanimous information sources such as titles,links,emails,pictures etc.The Web pages expressed by HTML,after analyzed by Internet Explorer,are suitable for browse,but not for machine processing as the way of data exchange.The principle of HTMLParser and related knowledge of regular expression,package HTMLParser and regular expression were introduced.Taking extracting email information inside websites as an example,the scheme of design was proposed.The principle of email extraction and key technique were presented.The algorithm of email extraction was given.URL extraction module,email extraction module and storage module were described in detail.The result of extraction is stored in database for the use of data retrieval.