Exploiting iteration-level parallelism in dataflow programs
Lubomir Bic, John M. A. Roy, Mark D. Nagel · 2003
An approach to extracting iteration-level parallelism from dataflow programs is presented. The method exploits the single-assignment principle, which guarantees that any data value has exactly one producer. To minimize interprocessor communication, the code is modified so that, at run time, each producer executes on the processor that holds the corresponding data. Overhead resulting from possibly remote read accesses is alleviated by a software technique similar to caching. The performance of this process-oriented dataflow system (PODS) is demonstrated using the hydrodynamics simulation benchmark called SIMPLE, in which a 19-fold speedup on a 32-processor architecture has been achieved.>