Modified Fuzzy K-mean clustering using MapReduce in Hadoop and cloud

Dweepna Garg, Parth Gohil, Khushboo Nirajkumar Trivedi · 2015

Apache Hadoop is an open source software framework which structures Big data (both structured and unstructured). It is nowadays one of the biggest motivator in market as data storage is inexpensive in it. The storage method of Hadoop uses a distributed file system which lets the user store large amount of data by simply adding more nodes to a Hadoop cluster. Clustering a large amount of data is a point of concern. MapReduce, a programming model used by Hadoop allows a parallelization technique by decomposing a larger problem involving large dataset to smaller portion of data and then executing it. A scalable machine learning library named as Mahout is an approach to clustering which runs on Hadoop. In this paper, the Hadoop multi-node cluster is formed using Amazon EC2. This paper focuses on Fuzzy k-mean clustering algorithm which is modified by centroid generation method using MapReduce in Hadoop. Experimental results depict a decrease in the number of iterations thereby leading to a decrease in the execution time when modification of Fuzzy K-mean clustering algorithm is done using Canopy generation in MapReduce in Hadoop.

Read the paper · More papers on PaperTik