Statistical Schema Integration across the Deep Web

Bin He, Kevin Chang · 2002

Schema integration is a central problem for integrating heterogeneous information sources. Traditionally, the problem has been defined and addressed as finding schema mapping between pairs of sources. This paper proposes a fundamentally different approach for schema integration, motivated by integrating large numbers of data sources on the Internet. On this ``deep Web, we observe two distinguishing characteristics that offer a fresh view for rethinking schema First, as the Web scales, there are ample sources that provide structured information in the same domains (e.g., books and automobiles). Second, while sources proliferate, their aggregate schema vocabulary tends to converge at a relatively small size. Motivated by these observations, we propose a novel paradigm, schema integration: Unlike traditional pairwise mapping, we take a holistic approach to integrate all input schema instances by finding an underlying generative schema model. We define a general statistical framework MGS for such hidden model discovery, consisting of hypothesis modeling, generation, and selection. Guided by the framework, we develop Algorithm MGSac for finding attribute correspondence. We demonstrate our approach over hundreds of real Web sources in four domains and the results show remarkable accuracy.

Read the paper · More papers on PaperTik