Framework for Distributed Semantic Web Crawler

Naresh Kumar, Manjeet Singh · 2015

Relevant information retrieval from the www mainly depends on the technique and efficiency of a crawler. So crawlers must be capable enough to understand the text and context of a link which they are going to crawl. Anchor text contains a very useful information to know about the target web page. Because knowledge about the target web page content helps the crawlers to decide their preferences of crawling the particular page. In this paper we have presented a design of distributed semantic web crawler capable of crawling both HTML and semantic web pages written using owl/RDf. In our crawler a component called page analyser is used to understand the theme of content of page and context of anchor tag in the page. The output of the page analyser is used to make crawling decisions. Our approach have revealed the great improvement in extracting the information from the links and guide the crawler for more relevant domain specific crawling.

Read the paper · More papers on PaperTik