Towards Abstraction of Heterogeneous Accelerators for HPC/AI Tasks in the Cloud
Antonio Maciá-Lillo, Higinio Mora, Antonio Manuel Jimeno-Morenilla, Tamai Ramírez-Gordillo · 2024
In the context of modern cloud computing, the integration of specialized hardware accelerators, including Graphics Processing Units (GPUs), Tensor Processing Units (TPUs), and Field-Programmable Gate Arrays (FPGAs), is pivotal for achieving high-performance computing (HPC) and supporting the intensive computational demands of artificial intelligence (AI) and machine learning (ML) workloads. Extensive work has been done in virtualization technology to support the use of these devices in cloud architectures. However, current research is focused on concrete technologies and devices. HPC and AI work benefit from the use of different accelerator devices. There is the need to generalize and account for the heterogeneity of devices and vendors. This paper proposes a conceptual model for a heterogeneous cloud architecture that schedules HPC and AI tasks across various accelerator devices. We propose the “Running Rounds” (RR) indicator, an analytical measure used in previous work, generalized for multiple GPU programming technologies, adapted for Intel GPUs, to predict the performance of jobs on different sets of resources. The efficacy of the model is demonstrated through experiments that validate the capability of the RR indicator to approximate performance, although with noted discrepancies that highlight the need for further refinement. The study underscores the potential and challenges of heterogeneous cloud environments.