A compiler framework for loop nest software-pipelining

Guang R. Gao, Lori Pollock, Alban Douillet · 2006

While improving the performance of micro-processors, computer architects have recently reached a technology wall. Higher frequencies are not sustainable anymore. The high and expensive power consumption and the lack of performance improvement on those uniprocessors have lead chip manufacturer to instead provide multi-threading capabilities to their current processor line. The trend goes further with multi-threaded cellular architectures where a chip is composed of hundred of thread units interconnected by an on-chip network and showing impressive raw performance numbers. However, the problem of harnessing so much computational power has yet to be solved. Several issues such as thread synchronization and programmability still exist. This dissertation proposes an elegant method, named Single-dimension Software Pipelining (SSP) to address those issues for an important class of programming structures, especially in the scientific domain: loop nests, perfect and imperfect. This dissertation shows how loop nests can be software-pipelined on both uniprocessor architectures and cellular architectures. The method subsumes modulo-scheduling as a special case for single loops. The entire framework is explained and includes: the handling of multi-dimensional dependences, the loop selection, the kernel generation, the register pressure evaluation, the register allocation and the code generation for both cellular architecture and uniprocessor architectures with dedicated loop hardware support. The method was implemented in the Open64 compiler and tested on the Intel Itanium architecture and on the IBM Cyclops64 architecture. Results show that SSP schedules outperform modulo-scheduling schedules on uniprocessor architectures and efficiently use the computational power of the cellular architectures.

Read the paper · More papers on PaperTik