Performance analysis of MapReduce on OpenStack-based hadoop virtual cluster
Nazrul Muhaimin Ahmad, Asrul Hadi Yaacob, Anang Hudaya Muhamad Amin, Subarmaniam Kannan · 2014
With the emergence of big data phenomenon, MapReduce and Hadoop distributed processing infrastructure have been commonly applied for large-scale data analytics. Hadoop distributed filesystem (HDFS) usually being deployed on physical clusters. With the advent of cloud computing platform such as OpenStack, a number of works have been carried out in implementing Hadoop virtual cluster on cloud computing infrastructure. This paper presents a performance analysis of MapReduce implementations on OpenStack-based Hadoop virtual cluster. The results of the analysis show that the MapReduce implementations are performing in a scalable manner towards an increase in the size of the Hadoop virtual cluster being deployed.