Computational and memory analysis of Tegra SoCs
Andrew Milluzzi, Alan D. George, Herman X. Lam · 2016
Low-power, embedded, GPU System-on-Chip (SoC) devices provide outstanding computational performance, especially for compute-intensive tasks. While clusters of SoCs for High-Performance Embedded Computing (HPEC) are not new, the computational power of these supercomputers has long lacked the efficiency of their more traditional, High-Performance Computing (HPC) counterparts. With the advent of the Tegra K1 and X1, the efficiency debate is significantly more complex. These new devices can provide up to 500 GFLOPS of Float32 performance with a TDP of just 10 Watts. This paper investigates the current state of NVIDIA SoCs, comparing and contrasting their performance and power characteristics to high-end NVIDIA accelerators. In order to perform this analysis, we leverage two metrics, Computational Density (CD) and External Memory Bandwidth (EMB), to provide a first-order estimate and then normalize the results with Realizable Utilization (RU). RU is a measurement of device efficiency, comparing observed benchmarking results to the theoretical CD or EMB. From this analysis, we are able to uncover which computational kernels might show similar or improved performance on an HPEC cluster of NVIDIA Tegra SoCs. Based on our observations, an SoC cluster would be plausible for some power-constrained HPC applications. Furthermore, the observed performance improvement in Tegra devices suggests that a future GPU-SoC cluster could be an option for a wide range of applications.