A Proposed Architecture for Parallel HPC-based Resource Management System for Big Data Applications
Waleed A. Al Shehri, Maher Ali Khemakhem, Abdullah Ahmad Basuhail, Fathy Elbouraey Eassa · Advances in Science Technology and Engineering Systems Journal · 2019
Big data can be considered to be at the forefront of the present and future research activities.The volume of data needing to be processed is growing dramatically in both velocity and variety.In response, many big data technologies have emerged to tackle the challenges of collecting, processing and storing such large-scale datasets.Highperformance computing (HPC) is a technology that is used to perform computations as fast as possible.This is achieved by integrating heterogeneous hardware and crafting software and algorithms to exploit the parallelism provided by HPC.The performance capabilities afforded by HPC have made it an attractive environment for supporting scientific workflows and big data computing.This has led to a convergence of the HPC and big data fields.However, big data applications usually do not fully exploit the performance available in HPC clusters.This is so due to such applications being written in high-level programming languages and do not provide support for exploiting parallelism as do other parallel programming models.The objective of this research paper is to enhance the performance of big data applications on HPC clusters without sacrificing the power consumption of HPC.This can be achieved by building a parallel HPC-based Resource Management System to exploit the capabilities of HPC resources efficiently.