Comparative analysis of K means clustering using different scheduling algorithms
D S Kavya, Chaitra D Desai · 2016
Hadoop is an open-source framework developed by Apache software foundation. It works on large datasets, as the data is been stored and processed across cluster in distributed environment. It consist of two components - HDFS and MapReduce. HDFS is used for storing the data and MapReduce is used to process those stored datasets. In current era large amount of data is generated by many websites, the data generated is enormous which is simply stated as "big data". These datasets are then processed in Hadoop framework. As the processing is done parallely in Hadoop framework appropriate scheduling of the resources is required to process large datasets in-order to gain better performance. Mainly scheduling algorithms are used to minimize completion time of a parallel application. Among users the resource allocation will be guaranteed by schedulers. In this paper we study about the performance of various scheduling algorithms like fair, fifo, capacity of Hadoop platform.