Design and Implementation of Basic Educational Web Resources Gathering System
Xu Chaojun · 2011
This paper introduces a topic specific web crawling system, which gathers basic educational resources from the web, and indexes them for the purpose of basic educational users. Compared to other similar theme based crawling system, the crawler integrates fuzzy rule based algorithm and VSM text analysis technology together to predicting each URL's relevancy to basic education while parsing current downloaded page HTML code. So, the system need not to save and retrieve low relevant URLs, and improve the system's whole efficiency greatly.