A Framework for Incremental Hidden Web Crawler
Rosy Madaan, Ashutosh Dixit, Komal Kumar Bhatia · 2010
Abstract—Hidden Web’s broad and relevant coverage of dynamic and high quality contents coupled with the high change frequency of web pages poses a challenge for maintaining and fetching up-to-date information. For the purpose, it is required to verify whether a web page has been changed or not, which is another challenge. Therefore, a mechanism needs to be introduced for adjusting the time period between two successive revisits based on probability of updation of the web page. In this paper, architecture is being proposed that introduces a technique to continuously update/refresh the Hidden Web repository. submitting the form in order to obtain the response pages containing the results of the query. In order to download the Hidden Web contents from the WWW the crawler needs a mechanism for Search Interface interaction i.e. it should be able to download the search interfaces in order to automatically fill them and submit them to get the Hidden Web pages as shown in Fig. 1.