Schema extraction for semi-structured data

Kam‐Fai Wong, Qiuyue Wang · 2003

XML has gained popularity as the standard of data exchange on the Web. It bears striking similarity to semi-structured data. Both of them are self-describing and no predefined schemas are required to create or use such data. However schema is very useful. It describes the data, helps data integration and facilitates query evaluation and storage management. Automatic schema extraction is necessary when we have a set of data but no schema is available. In this thesis, we investigate the problem in a general framework. We propose several extraction algorithms based on an incremental conceptual clustering method and automaton theory. Focusing on exploiting schema information in query evaluation, we implement a prototype system of storing and querying semi-structured data with graph schemas in an RDBMS and evaluate the algorithms.

Read the paper · More papers on PaperTik