PartSOM: A Framework for Distributed Data Clustering Using SOM and K-Means
L. Flavius, Jose Alfredo F. Cost · InTech eBooks · 2010
Self-Organizing Maps 44For that reason, several algorithms have appear aimed at clustering dispersed data into several locations and summarizing results in a central unit, ensuring data security and confidentiality.This approach is known as distributed data clustering (DDC).Clustering algorithms can be based on a wide variety of theories and techniques, including graph theory, combinatorial search techniques, fuzzy set theory, artificial neural networks, and kernels techniques (Xu & Wunsch II, 2005).Artificial neural networks are an important computational tool, with strong inspiration neurobiological and widely used in the solution of complex problems, which cannot be handled with traditional algorithmic solutions (Haykin, 1999).Applications for neural networks include pattern recognition, signal analysis and processing, analysis tasks, diagnosis and prognostic, data classification and clustering.Competitive neural networks provide a family of algorithms used for data representation, visualization and clustering.Among the unsupervised neural network models, the selforganizing map (SOM) plays a major role.SOM features include information compression while trying to preserve the topological and metric relationship of the primary data space (Kohonen, 2001).The SOM network defines, via unsupervised learning, a mapping of a continuous p-dimensional space to a set of model vectors, or neurons, usually arranged as a 2-D array.This work proposes a novel strategy for cluster analysis in distributed databases using a recently proposed architecture, named partSOM, and typical clustering algorithms, such as SOM and K-Means.In this approach, the clustering algorithm is applied separately in each distributed dataset, relative to database vertical partitions, to obtain a representative subset of each local dataset.In the sequence, these representative subsets are sent to a central site, which performs a fusion of the partial results.Next, a clustering algorithm is applied again to obtain a final result.The main contribution of this paper is to show that, in situations where the volume of data is very large or when data privacy and security requirements impede consolidation at a single location, the results obtained with the application of this strategy justify its use.The remainder of the chapter is organized as follows: section 2 presents a brief bibliographical review about distributed data clustering algorithms and section 3 describes the main aspects of the SOM and K-Means algorithm.Section 4 presents the proposed strategy, detailing its operation and the advantages obtained with its use.Section 5 describes the methodology used in the experiments and section 6 present the results of the application of the proposed strategy for some datasets, comparing them with the results obtained with the traditional approaches.Finally, section 7 presents conclusions and the direction of future algorithm research. How to referenceIn order to correctly reference this scholarly work, feel free to copy and paste the following: Flavius L. Gorgonio and Jose