Design of Crawler Based on HTML Parser Information Extraction

Ping Yi · Microcomputer Information · 2009

Whether general search engine or vertical search engine, the design of web crawler is the core technology. In this article, a novel system of life-theme web crawler based on HTMLParser information extraction is thoroughly studied. In this system, a simulation searcher is designed for collecting the seed URL by analyzing tree structure of life-theme website, then, based on the discussion of HTMLParser information extraction, the target URL that relate to life-theme is extracted from the seed pages. Empirical studies show that the Precision=93.552% and the Recall=96.720%, proving its effectiveness and achieving requirements for general enterprise-level application of vertical search engine.

Read the paper · More papers on PaperTik