Making Holistic Schema Matching Robust: An Ensemble Framework with Sampling and Voting
Bin He, Kevin Chen–Chuan Chang · Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign) · 2004
With the prevalence of databases on the Web, \\emph{large scale} integration has become a pressing problem. As an essential task, \\emph{holistic schema matching} (i.e., discovering attribute correspondences among many schemas) has been actively studied recently. As a ``data mining" approach in nature, holistic schema matching, on one hand, benefits from the large scale of input schema data, while on the other hand, also suffers the problem of noises. Such noises often inevitably arise in the automatic extraction of schema data, which is mandatory in large scale integration. For holistic matching to be viable, it is thus essential to make it robust against noisy schemas. Toward this goal, we propose a novel ``ensemble" framework, which aggregates a multitude of base holistic matchers to achieve robustness, by exploiting statistical sampling and majority voting: To begin with, we observe that Web query interfaces possess two interesting characteristics: 1) ``redundancy of attributes"-- that schemas tend to share attributes, and 2) ``non- dominance of noises"-- that noisy schemas are relatively few. These observations inspire us to develop a generic \\emph {ensemble} framework, which consists of \\emph{multiple sampling}, \\emph{ranking aggregation} and \\emph{matching selection}. In essence, our approach creates an ensemble of base holistic matchers, by randomizing the schema data into many \\emph{trials} and aggregating their ranked results by taking majority voting. We provide analytic justification of the robustness of the ensemble. Empirically, our experiments show that the ``ensemblization" indeed significantly boosts the matching accuracy, over automatically extracted schema data.