Analyzing and revising data integration schemas to improve their matchability
Xiaoyong Chai, Mayssam Sayyadian, AnHai Doan, Arnon S. Rosenthal, Len Seligman · Proceedings of the VLDB Endowment · 2008
Data integration systems often provide a uniform query interface, called amediated schema, to a multitude of data sources. To answer user queries, such systems employ a set ofsemantic matchesbetween the mediated schema and the data-source schemas. Finding such matches is well known to be difficult. Hence much work has focused on developing semi-automatic techniques to efficiently find the matches. In this paper we consider the complementary problem ofimproving the mediated schema, to make finding such matches easier. Specifically, a mediated schemaSwill typically be matched with many source schemas. Thus,can the developer of S analyze and revise S in a way that preserves S's semantics, and yet makes it easier to match with in the future? In this paper we provide an affirmative answer to the above question, and outline a promising solution direction, calledmSeer. Given a mediated schemaSand a matching toolM,mSeerfirst computes a matchability score that quantifies how wellScan be matched against usingM. Next,mSeeruses this score to generate a matchability report that identifies the problems in matchingS.Finally,mSeeraddresses these problems by automatically suggesting changes toS(e.g., renaming an attribute, reformatting data values, etc.) that it believes will preserve the semantics ofSand yet make it more amenable to matching. We present extensive experiments over several real-world domains that demonstrate the promise of the proposed approach.