Data Parallel Code Generation for Arbitrarily Tiled Loop Nests

Georgios Goumas, Nikolaos Drosinos, Maria Athanasaki, Nectarios Koziris · 2002

Tiling or supernode transformation is extensively discussed as a loop transformation to efficiently execute nested loops onto distributed memory machines. In addition, a lot of work has been done concerning the selection of a communication-minimal and a scheduling-optimal tiling transformation. However, no complete approach has been presented in terms of implementation for non-rectangularly tiled iteration spaces. Code generation in this case can be extremely complex, while parallelization issues such as data distribution and communication are far from being straightforward. In this paper, we propose a complete method to efficiently generate data parallel code for arbitrarily tiled iteration spaces. We assign chains of neighboring tiles to the same processor. Experimental results show that nonrectangular tiling allows better scheduling schemes, thus achieving less overall execution time.

Read the paper · More papers on PaperTik