Converting data-parallelism to task-parallelism by rewrites: purely functional programs across multiple GPUs
Bo Joel Svensson, Michael Vollmer, Eric Holk, Trevor L. McDonell, Ryan Newton · 2015
High-level domain-specific languages for array processing on the GPU are increasingly common, but they typically only run on a single GPU. As computational power is distributed across more devices, languages must target multiple devices simultaneously. To this end, we present a compositional translation that fissions data-parallel programs in the Accelerate language, allowing subsequent compiler and runtime stages to map computations onto multiple devices for improved performance---even programs that begin as a single data-parallel kernel.