Focused Crawler Based on Improved Algorithm of Web Content Similarity

Xiangwen Liao · Computer and Modernization · 2011

Focused crawler is an important part of the vertical search engine.The Web content relevance algorithm of traditional focused crawler only considers term frequency,ignores the location information of key terms.After the analysis of the focused crawler based on the Web content relevance,this paper proposes an improved method of calculating relevance using the features of HTML tags.Experimental results show that the average accuracy of improved algorithm is 64.99% and increases 15.37% compared to the original method.

Read the paper · More papers on PaperTik