Multi-Data Source Fusion in PDMS
Gilles Nachouki · 2008
Abstract. In this paper, we propose a new approach for data fusion in the context of schema-based Peer-To-Peer (P2P) systems. Schema-based systems manage and provide query capabilities for (semi-)structured information: queries have to be formulated in terms of schema. The dominating schema-based systems are relational databases or XML documents. Schema-based P2P are called Peer Data Management system (PDMS). A challenging problem in a schema-based peer-to-peer (P2P) system is how to locate peers that are relevant with respect to a given query. Our proposal, lies in application of the multi-data source fusion approach [17] in the context of PDMS. In this context, a multi-data source schema describes static or active data sources (in addition of their conflicts) shared by peers which are semantically neighbors. Static data sources can be structured or semi-structured (e.g. relational databases or XML documents), whereas active sources are services (e.g. Java Applications, Web Services etc.). Multi-data source schemas, distributed, shared and maintained by peers, are the basis of a semantic overlay network. This network and the power of Multi-data source Fusion Language (MFL) [17] are exploited, in this paper, for efficient query propagation towards the relevant peers. The design of our proposed system is, an extension of MDSManager [16][17], called Peer Multi-Data source Management System (PMDMS). We give a performance evaluation of the semantic query routing with respect to important criteria such as precision, recall, response time and number of messages. We compare next this result with the routing algorithm of SenPeer [4], a PDMS developed recently with respect to an hybrid (i.e. super-peer/peer) approach where peers are connected to superpeers according to their semantic domains. Our proposal compared to an hybrid approach shows a better performance regarding the precision, the response time and the number of messages exchanged between peers. 1