Detailed Analysis and Optimization of Irregular-Shaped Matrix Multiplication on Multi-Core DSPs
Haotian Mo, Qinglin Wang, Linyu Liao, Biao Li, Lihua Chi, Jie Liu · 2024
Irregular-shaped General Matrix Multiplication (GEMM) is extensively used in diverse workloads, such as scientific simulation and deep learning. In response to energy efficiency constraints, low-power multi-core digital signal processors (DSPs) have emerged as a viable alternative architecture in HPC systems. This study examines the performance of existing GEMM implementations working on irregular-shaped matrices and observes sub-optimal outcomes due to deficiencies in memory optimization and core-level parallelism extraction. For multi-core DSPs in FT-M7032, a CPU-DSP heterogeneous processor for HPC, we introduce dspIMM - a new multi-core parallel implementation for irregular-shaped matrix multiplications. dspIMM incorporates a new loop ordering, stack-space optimization, multi-dimensional core-level parallelization, a communication-computation overlap implementation with data prefetch across loops, and blocking optimization. These optimizations effectively enhance memory access and the core-level parallelism in irregular-shaped GEMMs. Experimental results demonstrate that the proposed communication-computation overlap optimization achieves the highest performance improvement in dspIMM, and the average speedup of dspIMM over previous implementations finally achieves up to 3.34 times.