Information retrieval for probabilistic schema matching
Cong Xi-hui · Jisuanji gongcheng yu sheji · 2008
Distributed information systems tend to be highly heterogeneous,integrate different computer platforms,data storage formats,document models and schemas which structure the documents and the latter aspectrequires to transform data structured under one schema into data structured under a different schema.For these reason,a probabilistic framework is introduced,called PMap.Our approach gives a probabilistic interpretation of the prediction weights of the candidates,selects the rule set with highest matching probability.Schema matching is the problem of finding correspondences(mapping rules,e.g.logical formulae) between heterogeneous schemas e.g.in the data exchange domain,or for distributed IR in federated digital libraries.The union formulae is formed by IR heterogeneous.