CoreTSAR: Task Scheduling for Accelerator-aware Runtimes
Tom Scogland, Wu-chun Feng, Barry Rountree, Bronis R. de Supinski · VTechWorks (Virginia Tech) · 2012
Heterogeneous supercomputers that incorporate computational ac-celerators such as GPUs are increasingly popular due to their high peak performance, energy efficiency and comparatively low cost. Unfortunately, the programming models and frameworks designed to extract performance from all computational units still lack the flexibility of their CPU-only counterparts. Accelerated OpenMP improves this situation by supporting natural migration of OpenMP code from CPUs to a GPU. However, these implementations cur-rently lose one of OpenMP’s best features, its flexibility: typical OpenMP applications can run on any number of CPUs. GPU imple-mentations do not transparently employ multiple GPUs on a node or a mix of GPUs and CPUs. To address these shortcomings, we present CoreTSAR, our runtime library for dynamically schedul-ing tasks across heterogeneous resources, and propose straightfor-ward extensions that incorporate this functionality into Accelerated OpenMP. We show that our approach can provide nearly linear speedup to four GPUs over only using CPUs or one GPU while increasing the overall flexibility of Accelerated OpenMP. 1.