Research and implementation for focused crawler based on probabilistic model

Liang Jiu-zhen · Computer Engineering and Science · 2013

Based on the study and research of the existing variety of focused crawlers,the paper proposes a focused crawler using probabilistic model,which analyzes various characteristics obtained in crawl process and uses probabilistic model to calculate each URL priority so as to filter and sort URLs.The proposed focused crawler based on probabilistic model solves the deficiency that most existing crawlers usually only adopt a single strategy for fetching webs from Internet.The distinct feature of our focused crawler is that:not only subject relativity but also history evaluation and web equality are considered so that the drift and tunneling problems are solved as well as the resource equality is guaranteed.Experimental results show that,compared with other focused crawlers,the focused crawler based on probabilistic prediction can gather more subject relevant web pages by retrieving less web pages,and has a better average topic relevant degree.

Read the paper · More papers on PaperTik