A configurable parallel iterative tuning framework
JinWei Zhao, Qi Zhu, Wei Wu, Hongmei Wei · 2021
Performance tuning is notoriously time consuming and laborious. Although a large number of compiling optimizations have been proposed, due to the complexity of the architecture of the current HPC system, it is still a challenge for people to make decisions which kind of code transforms should be performed. Iterative tuning is a mature method for parameter tuning. It could yield an exciting result with a price of many-times individual evaluation on different sets of compiling optimizations. However, in some scenarios, the price is so high that the iterative method is impracticable. In this paper, a configurable parallel iterative tuning framework is proposed to reduce the time cost of the iterative tuning method. With the help of the multicore parallelization, the pipeline between compilation and execution, programs are evaluated in a more effective way. A prototype system is built based on the TaihuLight supercomputer system. Evaluations on the micro-benchmarks and the real application show that the proposed framework is able to dramatically reduce the time cost of the iterative tuning method without degradation on optimizing capability.