Hybrid Algorithm on Semantic Web Crawler for Search Engine to Improve Memory Space and Time

Poonam Lambhate, Aparna Hambarde, M. Emmanuel, Shailesh Hambarde · 2021

World Wide Web is a colossal reservoir of hyperlink documents, these hyperlinks documents lay the foundation for communication on this omnipresent computing world. An acute need has arisen to develop and modify or design search algorithms that helps in efficiently and competently searching the specific required data from the huge repository available. A variety of search engines employ diverse web crawlers for obtaining search results efficiently. A variety of search engines use diverse web crawlers for obtaining search results proficiently. Some search engines use focused web crawler that collects different web pages that usually satisfy some specific property, by effectively prioritizing the crawler frontier and managing the exploration process for hyperlink. In this paper the main objective of this research is identifying the bottlenecks in the conventional framework. It conceives the experiment architecture which will showcase an improved technique to speed up the crawling process. This will lay the foundation for a future generation of current web. It will expand the horizon of crawling to diversify its application to the industry specific crawling mechanism. Shannon gain algorithm is used to determine the Threshold value of dynamic dataset. Experiments have been conducted on Apriori, Eclat, Declat Algorithms, Proposed hybrid algorithm. The comparative assessment of memory usage reflects the minimal consumption by the hybrid architecture. One of the main features of our proposed Method is its ability to tunnel through pages with a low score.

Read the paper · More papers on PaperTik