Accelerating the LOBPCG Method on Sunway TaihuLight

Yu Tianyu, Yonghua Zhao, Lian Qing Zhao · 2019

LOBPCG is a numerical eigensolver which can be parallelly implemented. In this paper, methods of optimization suitable for Sunway TaihuLight are discussed cover the main computations of LOBPCG. Two-level parallel architecture is proposed and followed by the introduction of a communication-avoiding parallel algorithm of sparse matrix-vector product. The algorithm not only overlaps computation and MPI communication but also overlaps data transfer between main memory and the scratch memory on the core group. A data buffering strategy with automatic adjustment during the runtime is proposed. Then, a new implementation of dense matrix multiplication adapted to the Sunway architecture is proposed. The scalability results demonstrate our routines led to significant promotions. We also analyze the performance on the Hamilton matrix from the Density Matrix Renormalization Group (DMRG) method, which is widely used by the computational physicist.

Read the paper · More papers on PaperTik