Ecosystem at Large: Hadoop with Apache Bigtop

Benoy Anthony, Konstantin Boudnik, Cheryl Adams, Branky Shao, Cazen Lee, Kai Sasaki · 2016

This chapter presents materials to get a better grip of data processing stacks powered by the software that forms the foundation of all modern Apache Hadoop distributions. Probably the most widely known data processing stack is Apache Hadoop™, composed of Hadoop Distributed File System (HDFS) and MapReduce (MR). Apache Bigtop is employed by all commercial vendors of Apache Hadoop as the base framework for their distributions. Conceptually, Bigtop is a combination of four subsystems serving different purposes: Bill of Materials (BOM), Gradle build system, iTest, and cluster deployment. Cluster deployment is implemented via Puppet recipes. In order to start work with Bigtop, user need to install JDK7 or later. Bigtop is currently in the almost completed transition from the Maven build system to Gradle. This proved to be more comprehensive and a better fit for the variety of the tasks needed to be managed in the Bigtop project.

Read the paper · More papers on PaperTik