Maintaining the repository of search engine freshness using mobile crawler

Mohammad Abu Kausar, Mohammad Nasar, Sanjeev Kumar Singh · 2013

The World Wide Web is a large repository of text documents, images, multimedia and much other information, referred to as information resources. A huge amount of fresh information is posted on the World Wide Web day by day. Web crawlers are programs that traverse the Web and download web documents in an automated manner. Search engines have to keep an up to date image to all Web pages and other web resources hosted in web servers in their index and data repositories, to provide improved and exact result to its users. The crawlers of these search engines have to retrieve the pages continuously to keep the index up to date. The list of URL is very huge and thus it is very difficult to refresh it quickly as 40% of web pages change daily. Due to which more of the network resources particularly bandwidth are consumed by the web crawlers to keep the repository up to date. This paper deals with a system based on web crawler using mobile agent. The proposed approach uses Java Aglets for crawling the web pages. The major advantages of web crawler based on Mobile Agents are that the analysis part of the crawling process is done locally at the residence of data rather than inside the remote server. This considerably reduces network load and traffic which can improve the performance and efficiency of the crawling process.

Read the paper · More papers on PaperTik