High-Performance Computing for Big Data
Jawwad Ahmed Shamsi, Muhammad Ali Khojaye · 2021
Big data systems require massive computational power to unleash their true potential. High Performance Computing (HPC) enables big data systems to incorporate parallelism and achieve low latency through efficiently distributing the workload among different nodes. Many of these systems are based on the Single Instruction Multiple Data (SIMD) architecture, where identical instructions are executed on different data items. HPC systems can be built either using general hardware systems or Application Specific Integrated Circuit (ASIC) such as Graphic Processing Units (GPUs) and Tensor Processing Units (TPUs). Different programming environments are used for HPC systems. For instance, OpenMP and Message Passing Interface (MPI) are used for general hardware in shared and distributed memory platforms, respectively. Similarly, GPUs utilize specific languages such as CUDA and OpenCL, whereas TPU-based systems can be developed using TensorFlow. For HPC systems, high-speed connects such as InfiniBand (IB) and NVLink can be used to ensure low-latency data transfer in distributed memory systems. HPC systems can also utilize volunteer computing or supercomputing models for resources. The chapter explores HPC systems and assessed their requirements and capabilities for solving the complexity and challenges related to big data problems.