A distributed, scalable parallelization of fuzzy c-means algorithm
R Dhivya Bharathi, Shailaja C. Shirwaikar, Vilas S. Kharat · 2016
Distributed Applications from different domains like Health care, E-Commerce, science, social networks etc., tend to generate large volumes of heterogeneous data that grow exponentially over a period of time leading to big data sets. Descriptive Analytics, on big data sets, pose a great challenge for traditional data analytical tools, since it is to be performed on the full data set, unlike predictive analytics which is done on training samples. Clustering is a commonly used descriptive analytics method, that requires necessary support for execution on big data sets. Clustering is a process of generating groups from a data set, such that members within a group have strong affinity to each other, and very weak affinity to members across the groups. The most common clustering algorithm is the K-means algorithm, which gives crisp clusters, where each data member belongs to exactly one cluster, This algorithm tends to provide less accuracy, when the nature of the data members is such that, they tend to show affinity towards more than one cluster. This vagueness in the data is better captured by the Fuzzy c-Means algorithm, which is seen as an improvement over the k-means, giving a set of fuzzy clusters, where the affinity to each of the clusters is defined by a membership value. In this paper, we present an implementation of fuzzy c-means algorithm, as a parallelized algorithm, for big data analytics, using the MapReduce programming model on Hadoop framework. A detailed performance analysis of the implementation, using various parameters of the algorithm is presented on a real data set from the machine library.