A review on techniques for optimizing web crawler results

Anuja Lawankar, Nikhil S. Mangrulkar · 2016

Now a days Internet is widely used by users to satisfy their information needs. In the exponential growth of web, searching for useful information has become more difficult. Web crawler helps to extract the relevant and irrelevant links from the web. To optimizing this irrelevant links various algorithms and technique are used. Discovering information by using web crawler have certain issues; such as different URLs having the similar text which increase the time complexity of the search, crawler resources are wasted in fetching duplicate pages and larger storage is also required to store these web pages. These are some of the roadblocks in getting optimum results from the crawler. This paper provides a deep study of existing information retrieval techniques (I.R) which would help researchers to retrieve optimum result links and information.

Read the paper · More papers on PaperTik