An Apotheosis Extraction Approach for Dynamic Web Data

Rajnish Kumar, Nidhi Bhirgu, Swati Shahi · 2013

Abstract — To extract the data from a hidden dynamic Web page is a very challenging work now a days because of existing many online databases. According to the survey, few years ago there were 10 million online databases but from the recent survey, now a days there are 30 million online database and the number is increasing day by day due to the revolution of technology in all the fields. Obviously, extracting the exact data is very hard. Already there are number of approaches which has been implemented but they generally filter the blocks after then they cluster, align and then extract the data but during this process, it is not guaranteed that we will get exact data. This Apotheosis Extraction Approach which is a fully experimented over number of online databases, automatically detect the schema of html, dhtml, jsp and all the webpages including scripting webpage.in this approach, when any query is submitted to web page then first of all, it form a tree in which all the part of html, dhtml, jsp is identified internally and it differentiate by positional, layout, appearance and content features after then clustering and regrouping of data, alignment of data is done. Sometimes we are able to get records but not items of data. For this purpose, we have generated a sleeve which will also be helpful during complex data extraction. This should not be happened. We should to get the proper data by using records and items. Records are nothing but the complete data and items are those parts which can be used for extracting. Suppose in above example, java 6 programming book is nothing but a record and ISBN, publication are the items.

Read the paper · More papers on PaperTik