High-performance programming frameworks for numerical simulation
Zeyao Mo · National Science Review · 2016
Over the past 20 years, a variety of research activities on high-performance programming have taken place worldwide. General-purpose application programming interfaces (GAPIs) have been designed and standardized; for example, MPI and OpenMP are two successful GAPIs in the field of numerical simulation. However, with the increase in performance of computer systems from teraflops to petaflops [1], these GAPIs will not be adequate to deal with emerging systems in the next 5–10 years. There are two main reasons for this [2]. First, GAPIs need to be significantly improved or redesigned to deal with the multilevel heterogeneous parallelism or deep memory hierarchies mandated by the power and microarchitecture considerations. Second, domain-specific abstractions for application programming interfaces (DAPIs) are essential to shield application developers from the complexity and availability of such GAPIs. To date, high-performance programming frameworks (HPPFs) have successfully been used in implementing DAPIs in the field of numerical simulation [2]. These frameworks can bridge the gap between application developers and GAPIs with the paradigm of ‘think in parallel and write sequentially’. Physics and numerical algorithms can be programmed sequentially based on a flowchart of the parallel code without any knowledge of GAPIs. Moreover, HPPFs allow parallel code to be ported from petaflop systems to emerging or exascale systems without having to be rewritten. In other words, HPPFs safeguard application software from changes due to evolving computer systems. Typically, HPPFs are in greater demand in institutes where there are more application domain programmers than high-performance programmers. Figure 1 depicts the inner workings of a HPPF. First, domain-specific data structures must be presented for the distribution of computation load toward potentially minimal data movement across the memory hierarchy for efficient implementation of numerical algorithms. Next, the domain-specific data dependencies of the numerical algorithms must be accurately described for the new data structures. Third, parallel computation models (PCMs) can be defined to depict different types of numerical computation phases. Here, each phase represents a type of data dependency among a set of sequential numerical computation entities. Fourth, parallel algorithms for data communication and load balancing must be designed for efficient implementation of the PCMs. Finally, DAPIs are defined as software components [3] using C++ constructs in which data structures are available and numerical computation entities are integrated. Then, the numerical algorithms can be implemented sequentially provided that they can be written as a flowchart comprising numerical computation phases for each existing type of PCM. Further, numerical algorithms can be classified into two types, namely, application aware and independent. The latter type can be clustered into numerical libraries, which can be integrated into HPPFs. Overview of an HPPF kernel. Possible software architectures for the construction of HPPFs are shown in Fig. 2. Six layers are required. The basic layer is a run-time system on top of a GAPI to schedule the various resources including processors, memory, input and output, process communication transactions, thread bindings and restart mechanisms, amongst others. The basic layer prepares all computing resources required by the numerical simulation. The second layer handles the data structures. Usually, domain partitioning methods are used. The computational domain is initially partitioned into the same number of subdomains as the number of computer nodes used. Then, each subdomain is further partitioned into regions, the number of which is equal to that of CPUs contained in each node. Finally, each region is partitioned into a number of patches greater than that of cores integrated in each CPU. In fact, the four nested levels (domain-subdomains-regions-patches) of the data structure are uniquely matched to the four hierarchical levels (system-nodes-CPUs-cores) of the computer architecture such that each node or CPU has a subdomain or a region, and each core has one or more patches. Based on the data structure, the third and fourth layers comprise parallel algorithms for data communication and load balancing, respectively. The fifth layer consists of libraries of numerical algorithms independent of the application, such as matrix operations, sparse linear system solvers, eigenvalue system solvers, fast discrete transformations, fast multiple methods, partial differential equation solvers ad so on. The sixth layer consists of the DAPIs. Software architecture of an HPPF. The software architecture shown in Fig. 2 is also suitable for heterogeneous computing using accelerators such as graphics processing units (GPUs) or Intel Xeon Phi (MIC). For example, patches can be distributed to accelerators and data transfer can be included for data movement between nodes and accelerators. The greatest challenge is to shield the heterogeneous parallel programming within each patch. Stencil template languages, belonging to the sixth layer, are useful candidates for numerical simulations. In the last two decades, HPPFs have resulted in great progress in the field of numerical simulation. Based on the types of grids for discretizing the computational domain, HPPFs can be classified into three types: structured grid, unstructured grid and combinational geometry. Dubey et al. [4] surveyed six typical frameworks for structured grids: BoxLib, Chombo, Cactus, Enzo, FLASH and Uintah. All have been developed over five years and have been used by third parties. SAMRAI, JASMIN, Overture, AMRClaw and PARAMESH are also summarized. The authors conclude that BoxLib, Chombo, JASMIN and SAMRAI are the most general for a wide range of applications and functionality. HPPFs for unstructured grids are widely used for finite element or finite volume computation of mechanism analysis. SIERRA [5] is one of the most general; UG and PHG are designed for finite element computation, and JAUMIN is similar to SIERRA. HPPFs for combinational geometry mainly focus on numerical simulation using Monte Carlo methods, of which GEANT4, MERCURY, OpenMC and JCOGIN are four examples. Mo surveyed the development of HPPFs [6] after the introduction of JASMIN, JAUMIN, JCOGIN and PHG in China with specific focus on many petaflop applications. Currently, high-performance computer systems are tending toward exascale computing. GAPIs are changing with respect to both power and performance efficiency. HPPFs also need to be rebuilt. For example, more partitioning levels may be needed in the data structures, which in turn require reconfiguration of the layers above, while many new scalable parallel algorithms need to be designed. However, the sixth layer for DAPIs should remain unchanged for fast migration of parallel code to emerging or exascale systems. Looking to the next 5–10 years, we anticipate two challenges for HPPFs: performance innovation toward exascale systems and safeguarding existing application developments. This work is under the auspices of National Science Foundation (91430218) and National High Technology Research and Development Program of China (2012AA01A309).