Build Energy-Efficient GPU Computing Environment for Machine Learning Algorithms with Register File Packing Technique
Xin Wang, Wei Zhang · 2023
Popular machine learning algorithms built with a mass of matrix multiplications can be well paralleled and the GPUs are desirable computing environment for these applications. However, the energy consumption on GPUs becomes a big concern which prevents the further performance increases for machine learning algorithms. In this work, we aim to build an energy-efficient GPU computing environment for famous machine learning algorithms with a GPU register file management theory named narrow width operand packing. First, we observed that RF occupancies of modern machine learning algorithms are relatively low leaving a great waste of GPU's RF leakage energy. Second, we found that the data maintained by the RF contains a large fraction of narrow-width operands for machine learning algorithms. We proposed to pack multiple narrow width operands to a single register. After the register packing, the RF occupancies can be further reduced. Finally, we attempt to save both static and dynamic energy consumption of GPU's RF by smartly shutting down the unused portion of the RF. We evaluated the energy reduction of this GPU RF management with five state-of-the-art machine learning algorithms. The experimental results show that the register packing techniques achieve the total GPU energy consumption reduction, up to 14.14% and 10.71 % on average.