Performance optimization of scientific applications on emerging architectures

Aiichiro Nakano, Hikmet Dursun · 2012

The shift to many-core architecture design paradigm in computer market has provided unprecedented computational capabilities. This also marks the end of the free-ride era—scientific software must now evolve with new chips. Hence, it is of great importance to develop large legacy-code optimization frameworks to achieve an optimal system architecture-algorithm mapping that maximizes processor utilization and thereby achieves higher application performance. To address this challenge, this thesis studies and develops scalable algorithms for leveraging many-core resources optimally to improve the performance of massively parallel scientific applications. This work presents a systematic approach to optimize scientific codes on emerging architectures, which consists of three major steps: (1) Develop a performance profiling framework to identify application performance bottlenecks on clusters of emerging architectures; (2) explore common algorithmic kernels in a suite of real world scientific applications and develop performance tuning strategies to provide insight into how to maximally utilize underlying hardware; and (3) unify experience in performance optimization to develop a top-down optimization framework for the optimization of scientific applications on emerging high-performance computing platforms. This thesis makes the following contributions. First, we have designed and implemented a performance analysis methodology for Cell-accelerated clusters. Two parallel scientific applications—lattice Boltzmann (LB) flow simulation and atomistic molecular dynamics (MD) simulation—are analyzed and valuable performance insights are gained on a Cell processor based PlayStation3 cluster as well as a hybrid Opteron+Cell based cluster similar to the design of Roadrunner—the first petaflop supercomputer of the world. Second, we have developed a novel parallelization framework for finite-difference time-domain applications. The approach is validated in a seismic-wave propagation simulation code on BlueGene/L, BlueGene/P and x86 quad-core processor based clusters. In addition, we have developed strategies for in-core optimization of the algorithmic kernel of this application, which is a high-order stencil computation—a common kernel to a spectrum of finite-differences based applications. Third, we have applied this systematic approach to a production level first-principles molecular-dynamics application, which has achieved a record of 2.58×10 12 electronic degrees of freedom on 163,840 BlueGene/P processors. Finally, we have devised a systematic end-to-end performance optimization scheme for large-scale scientific applications on emerging high-performance computing platforms.

Read the paper · More papers on PaperTik