Understanding Ultra-Scale Application Communication Requirements - eScholarship

Kamil Shoaib, John M. Shalf, Leonid Oliker, David E. Skinner · 2005

Understanding Ultra-Scale Application Communication Requirements Shoaib Kamil, John Shalf, Leonid Oliker, David Skinner CRD/NERSC, Lawrence Berkeley National Laboratory, Berkeley, CA 94720 ABSTRACT As thermal constraints reduce the pace of CPU performance improvements, the cost and scalability of future HPC archi- tectures will be increasingly dominated by the interconnect. In this work we perform an in-depth study of the commu- nication requirements across a broad spectrum of impor- tant scientific applications, whose computational methods include: finite-difference, lattice-bolzmann, particle in cell, sparse linear algebra, particle mesh ewald, and FFT-based solvers. We use the IPM (integrated Performance Moni- toring) profiling framework to collect detailed statistics on communication topology and message volume with minimal impact to code performance. By characterizing the paral- lelism and communication requirements of such a diverse set of applications, we hope to guide architectural choices for the design and implementation of interconnects for future HPC systems. risks limiting the extent of future scientific computation. In this work, we begin by examining the current state of HPC interconnects in the next section, motivating some of the characteristics we examine. Then we outline our pro- filing methodology, introducing the IPM library as well as the applications we study. We then explore the observed characteristics of our applications, including buffer sizes and connectivity, ending with an analysis of the feasibility of lower-degree interconnects for ultra-scale computing. INTERCONNECTS FOR HPC INTRODUCTION As the field of scientific computing matures, the demands for computational resources are growing at a rapid rate. It is estimated that by the end of this decade, numerous mission- critical applications will have computational requirements that are at least two orders of magnitude larger than cur- rent levels [1]. However, as the pace of processor clock rate improvements continues to slow, the path towards realizing peta-scale computing is increasingly dependent on scaling up the number of processors to unprecedented levels. As a result, there is a critical need to understand the range of CPU interconnectivity required by parallel scien- tific codes. The massive parallelism required for ultra-scale computing presents two competing risks. If interconnects do not provide sufficient connectivity and data rates for the communication inherent in HPC algorithms, application performance will not scale. If interconnects are over-built, however, they will dominate the overall cost and thus the extent of ultra-scale systems. Both performance scalability and cost scalability are crucial to manufacturers and users of large scale parallel computers. Failure to characterize and understand the communication requirements of HPC codes Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. To copy otherwise, to republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. IISWC 2005, Austin, TX, USA HPC systems implementing fully-connected networks (FCNs) such as fat-trees and crossbars have proven popular due to their excellent bisection bandwidth and ease of application mapping for arbitrary communication topologies. However, it is becoming increasingly difficult and expensive to main- tain these types of interconnects, since the cost of an FCN infrastructure composed of packet switches grows superlin- early with the number of nodes in the system. As supercom- puting systems with tens or even hundreds of thousands of processors begin to emerge, FCNs will quickly become in- feasibly expensive. Recent studies of application communi- cation requirements, such as the series of papers from Vetter and Mueller [18, 19], have observed that applications with the best scaling efficiency have communication topology re- quirements that are far less than the total connectivity pro- vided by FCN networks. Concerns about the cost and complexity of interconnec- tion networks on next-generation MPPs has caused a re- newed interest in networks with a lower topological degree, such as mesh and torus interconnects (like those used in IBM BlueGene/L, Cray RedStorm, and Cray X1), whose costs rise linearly with system scale. The most significant concern is that lower-degree interconnects may not provide suitable performance for all flavors of scientific algorithms. However, even if the network offers a topological degree that is greater than or equal to the application’s topological degree of connectivity (TDC), traditional low-degree interconnect approaches have several significant limitations. For exam- ple, there is no guarantee the interconnect and application communication topologies are isomorphic—hence prevent- ing the communication graph from being properly embed- ded into the fixed interconnect topology. Before moving to a radically different interconnect solution, it is essential to un- derstand scientific application communication requirements across a broad spectrum of numerical methods. In subsequent sections of this work we present a detailed analysis of application communication requirements for a

Read the paper · More papers on PaperTik