Efficient Execution of Dataflows on Parallel and Heterogeneous Environments
George Teodoro · Advances in systems analysis, software engineering, and high performance computing book series · 2012
Current advances in computer architectures have transformed clusters, traditional high performance distributed platforms, into hierarchical environments, where each node has multiple heterogeneous processing units including accelerators such as GPUs. Although parallel heterogeneous environments are becoming common, their efficient use is still an open problem. Current tools for development of parallel applications are mainly concerning with the exclusive use of accelerators, while it is argued that the adequate coordination of heterogeneous computing cores can significantly improve performance. The approach taken in this chapter to efficiently use such environments, which is experimented in the context of replicated dataflow applications, consists of scheduling processing tasks according to their characteristics and to the processors specificities. Thus, we can better utilize the available hardware as we try to execute each task into the best-suited processor for it. The proposed approach has been evaluated using two applications, for which there were previously available CPU-only and GPU-only implementations. The experimental results show that using both devices simultaneously can improve the performance significantly; moreover, the proposed method doubled the performance of a demand driven approach that utilizes both CPU and GPU, on the two applications in several scenarios.