Research on Web Crawler Module of Search Engine
XU Dong-ping · Modern Computer · 2010
Nowadays Internet resources expand rapidly,search engine extracts a clear path from a broad array of clutter information,so users can access the information as they need.Web crawler module which is implemented by Web spider program,is the basis of search engine,and makes or mars the search engine from the angle of resources.In view of the above,introduces the elementary principles of search engine,analyzes the work process of Web crawler module,after that studies the key components of open source Web spider Heritrix.Extends the extractor component on these bases,and achieves personalized crawling logic.