D5.2: Market and Technology Watch Report Year2
Aris Sotiropoulos · Zenodo (CERN European Organization for Nuclear Research) · 2019
This document is the second deliverable of PRACE-5IP Work Package 5 “Task 5.1 - Technology and market watch” and represents a periodic annual update on technology and market trends. It is thus the continuation of a well-established effort to carry out an assessment of the HPC market based on market surveys, supercomputing conferences, and exchanges between vendors and experts involved in the work package. The Top500 (Nov 2018) list is still dominated by China and the United States, that rank very close to each other, with, this time, the United States in the first and second place, with 21.8% of system share and 37.7% of performance share, China on third and fourth position, with 45.4% of system share and 31% of performance share. Two machines installed in Europe are in the Top10: SuperMUC-NG (LRZ, Germany) and Piz Daint (CSCS, Switzerland). Japan still dominates the Green500 list with four supercomputers ranked in the Top10. The first and second ranked supercomputers in the Top500 list, Summit and Sierra, are also the machines showing the best HPCG performance. The only European supercomputer in the Top10 of HPCG is Piz Daint. A new benchmark related to IO performance (IO500) initiated in 2017 is gaining momentum with 63 entries while a proposal for a new benchmark on large scale deep learning called Deep500 have been presented. Plans for exascale are well-defined in China, USA and Japan. In China, three tracks are explored in parallel with three prototypes deployed by Sugon, Tianhe, and Sunway (ShenWei) in 2018. In Japan, the post-K project, targeting exascale class supercomputer in 2021, is progressing on schedule with the test started on the first prototype of the ARM-based CPU developed by Fujitsu. In the United States, the ECP project is addressing application development, software technology, and hardware technology and exascale systems testbeds. In Europe, EuroHPC is a Joint Undertaking established in 2018 with, as one of its goals, to construct an exascale supercomputer based on European technology. HPC as Cloud computing offering has already been available for some time and is a real option for some HPC workloads. It generally complements, rather than replaces, the traditional HPC systems. Big players like Amazon, Microsoft and Google have commercial offerings that target the HPC market. The main trend in 2018 for HPC is bare metal service, while solutions based on containers have some advantages in terms of redundancy and enhanced information security. Offers including GPU, targeting mostly AI, may also be of interest for HPC. Even quantum computing (IBM Quantum System for example) is also available in the cloud. Regarding the consolidation in the HPC market, the acquisition of Cavium by Marvell is an important event since the ThunderX ARM-based processors are among the most promising for HPC. In the EU landscape, EuroHPC is the central and overarching new piece, consolidating or recomposing the ecosystem. It brings the promise of better coordination and more funding for HPC in Europe, both for infrastructures and for R&I, including applications. In this context, existing entities like PRACE, ETP4HPC, and BDVA, need to evolve in order to take into account this new context. Among the coordination and support actions, EXDCI-2 continues the coordination of the HPC ecosystem with important enhancements with respect to EXDCI, so as to better address the convergence of big data, cloud and HPC. On the R&I side, centers of excellence and FET HPC projects remain important and active players in the field while the new EPI (European Processor Initiative) is of major importance for Europe. On the infrastructure side, PPI4HPC and HBP are contributing respectively for deploying innovative equipment in HPC centers and developing a data infrastructure complementing the current supercomputer infrastructure. In terms of core technologies for HPC system processors, Intel is now facing the competition of new competitors: AMD with the EPYC architecture, and Marvell, with the ThunderX ARM-based processor family. Several large systems based on these processors have already been announced. The ARM ecosystem is very active with processors targeting HPC announced by Huawei and Fujitsu. The POWER processor, coupled with NVIDIA GPU, is used for the supercomputers Summit and Sierra, which are listed on position #1 and #2 of the Top500 list as of November 2018. In terms of accelerators for HPC, these systems mostly rely on NVIDIA GPUs with AMD being so far the only competitor but with less success than NVIDIA. FPGA is still a niche market. The most used volatile memory is currently DDR4 which tops with a DIMM size of 128GB running at 2933MT/s. Another class of volatile memory technology important to HPC is High Bandwidth Memory (HBM), today usually implemented as stacked memory, allowing parallel access to multiple (up to 8) slices. It supports speeds up to 2GT/s. Non-volatile memory technologies are still emerging but may be used in the future on compute node and on IO systems. Regarding data storage and data management, technologies and components are evolving, new services are being deployed some targeting specific workload like AI. Tape storage is still widely used with capacity increasing over time. This is the same for disk storage, DDN being the main player. Flash and non-volatile memories are expected to play an increasing role in the near future. In terms of vendor solutions and roadmaps, HPE remains in the 2nd position in the HPC world with products derived from SGI (acquired by HPE). HPE has installed the largest ARM-based cluster in Sandia National Lab. Fujitsu continues to be the main provider of high-end HPC solutions in Japan. Chinese vendors (mostly Huawei, Lenovo) are actively investing in the HPC market. Atos/Bull currently has 22 supercomputers ranked in the Top500 and has announced a blade for ARM processors. Paradigm shifts in HPC technologies include AI, heterogenous architectures, and neuromorphic computing. Since our last report, AI is gaining momentum in a number of dimensions: to assist regular HPC simulations (“augmented HPC”) or to replace HPC simulation. HPC can also be used to create high fidelity training AI databases. Modular heterogeneous solutions are becoming widespread with the double objectives of limited power consumption and programmability. In this approach, the user will no longer see a supercomputer as a single homogeneous entity but rather as an aggregation of specialized modules sharing common resources such as parallel file systems or visualization nodes on a central network. The term “neuromorphic computing” broadly refers to compute architectures that are inspired by features of brains as found in nature. There is a strong interest, at the level of research, into such devices in the context of brain modelling as well as artificial intelligence and machine learning.