A genetic algorithm based focused web crawler for automatic webpage classification

Nishtha Goyal, R. Bhatia, Manish Kumar · 2016

The rapid increase in the amount of information present on the World Wide Web makes it difficult to find information of interest to a user. Search engines uses focused Web crawlers to get the information about a particular topic. Focused Web crawler seeks, gathers and maintains webpages relevant to a pre-defined set of topics rather than downloading all the webpages. During focused crawling, automatic webpage classification method is used to determine whether the webpage is on-topic or not. This paper discusses a genetic algorithm based automatic webpage classification technique. In this method, tags and terms are considered as features and the classifier is made to learn from the webpages in the training set. The best features are selected from the genetic algorithm based fitness optimization technique. Using both tags and terms as features, high precision on test data is achieved.

Read the paper · More papers on PaperTik