Parallelization of LU Decomposition on the Godson-Tv1 Many-Core Architecture
Long Guo · Chinese Journal of Computers · 2009
The many-core architecture is increasingly becoming a promising computing platform due to the advancement of semi-conductor technology. LU decomposition is a widely used kernel in both scientific and engineering computations. Although there are a lot of related works on traditional parallel architectures,there is still little work focusing on parallelizing it on many-core architectures. This paper investigates this problem from three aspects: load balancing,latency hiding and performance modeling. There are three contributions of this work: Firstly,a novel load balancing technique has been introduced to overcome the limitations of 2D scatter decomposition. Experimental results show that the proposed scheme achieves 20% performance improvement without optimization and 40% improvement after optimization. Secondly,an analytical performance model is presented. Quantitative experimental study shows that by carefully hiding memory latency through on chip memory hierarchy and for a selected block size,the upper bound of theoretical performance can be approximated by experiments. Experimental results also reveal two primary causes which make theoretical speedup hard to achieve: limited DRAM bandwidth and resource contention of on-chip network.