Use of Genetic Algorithm to Find Average Number of Generations to Converge for Web Retrieved Documents Using Jaccard Similarity Coefficient
Vivek Jaglan · 2013
The rapid growth of the world-wide web poses unprecedented scaling challenges for general-purpose crawlers and search engines. A focused crawler aims at selectively seek out pages that are relevant to a pre- defined set of topics. Besides specifying topics by some keywords, it is customary also to use some exemplary documents to compute the similarity of a given web document to the topic. In this paper we present a method for finding out the most relevant document for the given set of keyword by using the method of similarity measure of Jaccard coefficient. Due to the randomized nature of genetic algorithm we show that generation number for convergence value of 1 is not always the same. The similarity coefficient for a set of documents retrieved for a given query from Google are find out then average relevancy is calculated. In this paper we have averaged 10 different generations for each query by running the program 10 times for the fixed value of Probability of Crossover Pc=0.7 and Probability of Mutation Pm=0.10. The same experiment is conducted for 10 queries.