Memory-coherence between host and devices in a runtime
Rubén Cano, Daniel Jiménez-González, Vicenç Veltran · UPCommons institutional repository (Universitat Politècnica de Catalunya) · 2020
As the end of the Moore’s law approaches, more specific devices such as GPUs, FPGAs or AI accelerators tend to steal the workload that was traditionally run on the CPU, allowing with this offload more specific solutions that improve the execution time of specific applications. One of the main problems that arise with this approach, is that now, the data is not centralized in one main memory, but distributed among the different accelerators which need a correct and coherent data to perform its operations. This can potentially limit the performance an accelerator can achieve, as well as delegates the programmer the task of enforcing the coherence between memories. To relieve this model, in which the programmer has to take into account the memory of devices, models like NVIDIA Unified Memory[1] manage the hard work of maintaining the memory-coherence, potentially hurting performance but making the per-device memory management much easier. In this work, the main objective is to develop an extension for a task-based runtime, which maintains the coherence between SMP and multiple devices in the system, using the dependency information of a task, acting as a Translation- Allocation Layer between the multiple memory spaces defined by the accelerators. For our development, we use Nanos6 Runtime[2] as the target of our implementation. Nanos6 is a Runtime that implements the OmpSs-2 programming model, it is an SMP taskbased runtime, with cluster support and is able to run CUDA tasks using unified memory. With this implementation, our objective is to extend Nanos6 capabilities to allow running CUDA with distributed memory, as well as other kind of accelerators such as FPGAs.