Learning Capable Focused Crawler for Information Technology Domain
Mukesh Kumar, Renu Vig · International Journal of Computer Applications · 2012
ABSTRACT The Web provides us with a huge and endless resource for information. But, the rapidly growing size of the Web poses great challenge for general purpose crawlers and search engines. It is impossible for any search engine to index the whole Web. Focused crawler collects domain relevant pages from the Web by avoiding the irrelevant portion of the Web. Focused crawler can help the search engine to index all documents present on the Web related to a specific domain which in turn provides the search engine‟s users complete and up-to-date contents. In this paper we present a focused crawler capable of learning from the previous crawl results to collect the relevant documents. Crawling results for three consecutive learning phases are shown. Results indicate significant improvement in terms of relevancy to the focused domain. Keywords Web, Internet, Retrieval, Focused Web Crawler, Search Engine etc. 1. INTRODUCTION With the ongoing growth of web, finding the right information becomes an increasingly difficult task which often leads to undesired results. This made it important to develop document discovery mechanism. A crawler is a program used by search engine that retrieves Web pages by wandering around the Internet following one link to another. Web search engines such as Goggle, AtlaVista provides access to the Web documents. A search engine‟s crawler collects Web documents and periodically revisits the pages to update the index of the search engine. Due to the Web‟s huge size and dynamic nature, Ari Pirkola (2007