A Parallel Heterogeneous Computation Library for Embedded ARM Clusters
Nikhil Khatri, Priya Bagaria, T. S. B. Sudarshan · 2020
In recent years several hobbyists and research institutions have turned to ARM based cluster computers to meet their high performance computing needs. These systems are typically low cost, power efficient machines made of 4 or more nodes. However, current applications running on these machines are limited to the computational capabilities of the on-board CPUs. An attempt to fully utilise the available hardware is to use the on-board GPUs for general purpose computation. The challenge with this approach is the difficulty of programming for embedded ARM GPUs, which support a very limited programming interface such as OpenGL ES 2.0. In this paper, we present a system which allows a high level program to utilise multiple embedded ARM based GPUs in parallel to perform computation, without the application developer concerning themselves with the details of OpenGL and writing shaders. We describe the vector and matrix functions which we implement and provide results detailing their speedup. For matrix multiplication, our system shows a speedup of nearly 7X on the 8 nodes of the PHINEAS cluster computer. We further describe how these functions may be composited with some CPU functionality to implement inferencing for a Convolutional Neural Network trained in Keras on a different machine. Our benchmarks demonstrate that the system's performance scales well across several nodes for a variety of tasks. Our work brings the benefits of GPGPU computing to an as yet untouched class of cluster computers.