Structure based Data Extraction from Hidden Web Sources: A Review

Author Anuradha, Anukrati Sharma · International Journal of Computer Applications · 2011

In order to extract data from the web pages of Hidden web sources, many semi-automatic and automatic techniques are proposed based on structure and tags of HTML documents.These techniques include machine learning and schema-matching approaches to solve the problem of data extraction.This paper discusses the research that has been done in the area of data extraction from Hidden Web sources.The goal of this paper is to discuss the advantages and disadvantages of currently existing techniques.

Read the paper · More papers on PaperTik