The influence of data replication in the knowledge discovery in distributed databases process

Valentin Pupezescu, Radu Rădescu · 2016

The process of knowledge discovery applied in distributed databases implies finding useful knowledge from mining data sets stored in real implementations of distributed databases. Distributed Databases represents a software system that allows a multitude of applications to access the data stored in local or remote databases. In this scenario, the data distribution is achieved through the process of replication. Nowadays many solutions for storing the data are available: relational distributed Database Management Systems (DBMS), NoSQL storing solutions, NewSQL storing solutions, graph oriented databases, object oriented databases, object-relational databases, etc. The present study analyzes the most commonly used storing solution: the relational model. The replication topology used in the related experiments was the classical publisher-subscriber topology. The distribution of data is made from the publisher system. The present work studies the interaction between the most suited distributed data mining architecture (Distributed Committee Machines) for mining distributed data and real relational distributed databases. The chosen Data Mining task is the classification one. Distributed Committee Machines are a group of neuronal networks working in a distributed procedure to obtain an improved classification performance compared to a single neural structure. In these experiments we used the classical multilayer perceptron trained with the backpropagation algorithm. The execution performance of the Distributed Committee Machine is analyzed, based on some of the most used types of replication in relational databases: snapshot replication, merge replication, transactional replication, and transactional with queued updating replication. For all these types of replications the execution performances (distributed speedup and distributed efficiency) of the entire system is also analyzed. These results are useful to numerous research fields: adaptive e-learning applications, medical diagnosis, artificial intelligence, business management, etc.

Read the paper · More papers on PaperTik