Priority Queue Based Estimation of Importance of Web Pages for Web Crawlers

Mohammed Rashad Baker, Muhammet Ali Akcayol · International Journal of Computer and Electrical Engineering · 2017

There are hundreds of new web pages that are added daily to web directories.Web crawlers are developing over the same time of web pages growing up rapidly.Thus, the need for an efficient web crawler that deals with most of the web pages.Most of the web crawlers do not have the ability to visit and parse pages using URLs.In this study, a new web crawler algorithm has been developed using the priority queue.URLs, in crawled web pages, have been divided into inter domain links and intra domain links.The algorithm sets weight to these hyperlinks according to the type of links and stores links in the priority queue.Experimental results show that the developed algorithm gives a well crawled performance against unreached crawled web pages.In addition, the developed algorithm has a good capability to eliminate duplicated URLs.

Read the paper · More papers on PaperTik