Deep Web Integration: the Tip of the Iceberg

Yasser Saissi, Ahmed Zellou, Ali Idri · International Review on Computers and Software (IRECOS) · 2015

The web is divided in two parts, a part that search engines can access and which is called the surface web, and an inaccessible part called the deep web. The deep web is much bigger and richer in information than the surface web, and its web sources are only accessible through the associated Html forms. Our aim in this paper is to present our automatic approach to extract a relational schema describing a selected deep web source. This relational schema can be used by a virtual integration system to access the associated deep web source. Our approach is based on a static and dynamic analysis of the Html forms giving access to the selected deep web source. Our approach process uses two external knowledge databases: The first one is our proprietary knowledge database about the deep web domains called the Identification Tables and the second one is an external ontology. All the information extracted by our approach from and through the associated Html forms are used subsequently to build our final relational schema describing the associated deep web source.

Read the paper · More papers on PaperTik