Electron Dynamics Simulation with Time-Dependent Density Functional Theory on Large Scale Symmetric Mode Xeon Phi Cluster
Yuta Hirokawa, Taisuke Boku, Shunsuke A. Sato, Kazuhiro Yabana · 2016
Many-core architecture processors such as the Intel Xeon Phi provide new solutions for high performance computing (HPC) systems that require high computing performance combined with low power consumption. However, this computing efficiency is difficult to achieve due to its characteristics as the "throughput core" which differs from the ordinary "latency core" characteristic of advanced processors such as Intel Xeon. In this study, we implement a real scientific code named Ab-initio Real-time Electron Dynamics simulator (ARTED), which is an electron dynamics simulator on (Time-Dependent Density Functional Theory (TDDFT). ARTED runs on an Intel Xeon Phi cluster using the Symmetric mode, where all CPU and Intel Xeon Phi resources contribute to the computation. A kernel of stencil computation that dominates the total computation time is optimized in a single-thread (single-core) level with various techniques such as explicit vectorization with 512-bit SIMD instructions. Consequently, the Native mode operation of the Intel Xeon Phi achieves approximately twice the level of performance of the Intel E5-2670v2 CPU (Ivy-Bridge, 1 socket, 10 cores) for the stencil computation with double-precision complex value: 212.2 GFLOPS with Xeon Phi and 106.9 GFLOPS with CPU. Moreover, the entire code achieves 1.45 times better performance compared with the CPU. For the Symmetric mode operation, we need to balance the load among different types of processors. We control the number of processes in the wave space for each processor to balance their loads. Consequently, the Symmetric mode operation achieves 2.16 times better performance compared with the CPU-only utilization on each node with strong scaling.