Using a Joint Link Similarity Evaluation Based Method for Crawling the Resources on Web
Zhang Nai · Chinese Journal of Computers · 2010
For many fields of Web research,how to fetch the interesting resources is crucial.At present,the chief method for obtaining the domain-specific resources on Web is to adopt the strategy of focused crawling.However,for the most current techniques of focused crawling,there are many problems in simultaneously meeting the high efficient crawl and the high quality of crawl results.This paper proposes a joint link similarity evaluation based algorithm.When evaluating the similarity between a link and a specific topic,the algorithm combines the direct evidence with indirect evidence on the topic similarity of the link.The direct evidence can be obtained by computing the topic similarity of the anchor text corresponding to the link.As to the indirect evidence,this paper presents a Q learning based algorithm for incrementally learning Web Link graph.The algorithm firstly builds a Web link graph by exploiting the on-topic Web pages fetched by focused crawler and then gets the map relationship between the link and topic similarity through online learning.Modeling any link as a multi-attribute vector,the system gives the link evaluator the ability to map the current link into the space of the Web link graph and thus obtains its approximate topic similarity.The experimental results for three specific topics show that the algorithm can significantly improve the precision and the recall of crawl results.