V-Hadoop: Virtualized Hadoop using containers
Srihari Radhakrishnan, Bryan J. Muscedere, Khuzaima Daudjee · 2016
MapReduce is a popular programming model used to process large amounts of data by exploiting parallelism. Open-source implementations of MapReduce such as Hadoop are generally best suited for large, homogeneous clusters of commodity machines. However, many businesses cannot afford to invest in such infrastructure and others are reluctant to use cloud services due to data security and privacy concerns. In this paper, we present V-Hadoop, a framework that leverages Linux containers to allow users to run Hadoop jobs efficiently without requiring large, expensive, physical machine clusters. We describe our design and implementation of V-Hadoop and show that it can effectively support cluster-level parallelism. We experimentally demonstrate that V-Hadoop is a viable solution that performs competitively compared to solutions designed for large clusters.