Design of focused crawler based on feature extraction, classification and term extraction

Shilpi Gupta · International Conference on Computing for Sustainable Global Development · 2016

A web crawler is a software program that scans the hypertext layout of the web pages, starting from a set of seed pages. The crawler retrieves these pages, indexed them and extracts the hyperlinks inside these pages to find out the addresses for more pages to be crawled. There is a set of problems emerging in the Automatic publication data gatherer, this paper gives a solution to the problems. A proposed architecture is given to improve the performance of the crawler by using the different feature extraction technique and the classification technique and also use search engine queries.

Read the paper · More papers on PaperTik