Large data computing using Clustering algorithms based on Hadoop

Samrudhi Tabhane · 2015

Abstract — The Hadoop Distributed File System (HDFS) is designed to store large data sets reliably and to stream those data sets at high bandwidth to user applications. In a large cluster, thousands of servers and host are directly attached and execute user application tasks. By distributing storage and computation across many servers, the resource can grow with demand while remaining economical at every size. Hadoop is a popular opensource implementation of MapReduce for the analysis of large datasets. To manage storage resources across the cluster, Hadoop uses a distributed user-level file system. This paper analyzes the performance of two major clustering algorithms K-means and DBSCAN on Hadoop platform and uncovers several performance issues. The experimental result demonstrates that K-means clustering algorithm is more efficient than DBSCAN algorithm based on MapReduce. Experimental results also show that DBSCAN algorithm based on MapReduce alleviates the problem of time delay caused by large data sets.

Read the paper · More papers on PaperTik