Thoughts on high-performance computing
Xuejun Yang · National Science Review · 2014
Parallel computing is the main technical approach for achieving very high performance computing. In the history of parallel computing, there have been three phases, i.e. moderate parallelism described by Amdahl's law [1], large-scale parallelism described by Gustafson's law [2], and high-productivity parallelism described by the productivity evaluation model [3]. In April 2010, IBM Inc. in their report ‘Some Challenges on Road from Petascale to Exascale’ presented five challenges in an exascale system; these stem from power consumption, memory access, communication, reliability, and programming [4], respectively referred to as the energy wall, memory wall, communication wall, reliability wall, and programming wall. Here SR is the reliability speedup, R(P) reflects the relation between fault-tolerant overhead and the number of nodes P, U is the traditional efficiency of the system, and UR is the reliability efficiency after introducing a reliability factor. Thus, the reliability wall is defined as the supremum of reliability speedup [5], which is a trade-off between reliability and computing performance when the system size is scaled up, as shown in Fig. 1. Similarly, memory, communication, and other walls are also defined as the supremum of the corresponding speedups. Reliability speedup and reliability wall [5]. Faced with the challenges of ‘walls’, we have made breakthroughs in both the architecture and enabling technologies. With respect to the architecture, we highlight (1) a design that balances computing with communication and memory access, coordinates the application, software, and hardware, and integrates the computer system and execution environment; (2) a chip architecture that integrates a general processor as well as a specific processor; (3) a support framework for applications; (4) state-of-the-art cooling technology, such as the ICEOTOPE Corp. cooling product; (5) a system-on-chip architecture based on neural networks; (6) a new programming language together with its compiler; (7) large-scale parallel algorithms that can be scaled up to tens of thousands or even millions of nodes; (8) a domain-oriented support environment for high-performance computing; (9) a support environment for the design and execution of parallel programs at the instruction, thread, multi-core, and multi-node levels; (10) the optimization of memory access at the architecture, operating system, compiler, and algorithm levels; (11) power optimization at the chip, architecture, operating system, and compiler levels; (12) fault tolerance technology integrated both the software and hardware; and (13) service-oriented cloud computing, amongst others. With respect to the enabling technology, with the rapid development of nanomaterial, quantum computing, and bioinformatics, we focus on (1) quantum walk Boson sampling, (2) programmable nanometer circuits, (3) memristors, (4) holographic optical storage, (5) on-chip optical interconnects, (6) system-wide optical interconnects, and so on. Information science and information technology complement one another. In the late 20th century, information technology developed rapidly. However, there has been no major breakthrough in the foundation of information technology, i.e. information science, in the last 40 years. Therefore, we need to strengthen research on the fundamental theory by refining scientific problems from engineering, enabling breakthroughs in the theory, and then applying these to engineering. We also emphasize interdisciplinary studies by promoting the intersection and merging of computer science with several disciplines, including physics, mathematics, material science, chemistry, and micro-electronics, in order to enhance the capacity for sustainable development.