Performance improvement for applications on parallel computers
Xiangzhen Qiao · 2005
The performance of applications on parallel computer systems is analyzed, and a general timing formula is proposed in the paper. The analyses are based on the double-stream (computation and data access) idea. According to these analyses, some techniques for improving the performance of applications on parallel systems are given. The cache thrashing phenomena on symmetric multiprocessors (SMP) are analyzed, and the fix point iteration method is used to deal with some cache thrashing phenomena. The timing balance problem of heterogeneous clusters is discussed, and the non-uniform partition strategy is studied. Some optimization techniques (horizontal-vertical decomposition, non-uniform partition, and overlapping of communication with computation) are used to improve the performance of a class of parallel applications on clusters of multiprocessors. Compared with previous results, our experiences show distinct performance improvements.