A Transfer-Aware Runtime System for Heterogeneous Asynchronous Parallel Execution

Soukaina N. Hmid, José G. F. Coutinho, Wayne W. Luk · ACM SIGARCH Computer Architecture News · 2016

This paper presents a novel resource management approach for efficiently managing the computation and the data movements between the host and its accelerators in a heterogeneous platform. Our approach is based on OmpSs, with support for multi-core CPUs, GPGPUs and Maxeler Data Flow Engines based on FPGA technology; it exploits data locality, data transfer costs and data dependencies. The proposed approach is supported by an offline learning process coupled with online monitoring, allowing performance to be estimated while learning from past observations during execution. Its performance is compared against the current OmpSs scheduler using five benchmarks: matrix multiplication, bitonic sort, N-body simulation, Cholesky decomposition and AdPredictor. The results show the proposed approach can achieve up to 4.25 times speed-up for Cholesky decomposition. Moreover, an evaluation with AdPredictor indicates that the FPGA version is up to 46 times faster than the CPU version for large task sizes.

Read the paper · More papers on PaperTik