Low Precision and Efficient Programming Languages for Sustainable AI: Final Report for the Summer Project of 2024
Joao Vitor De Oliveira Silva, Tokey Tahmid, Weslley da Silva Pereira · 2024
This document contains all relevant material generated during the authors' summer internship at NREL in 2024.This report shows how to improve the energy efficiency of a few code samples by using low-precision data types combined with mixed-precision algorithms.The main applications considered here are (i) linear system solvers using mixed precision, and (ii) neural networks using mixed precision.This report also discusses how programming languages affect the energy consumption of algorithms, energy metrics for a code and energy measurement tools, and the currently available software and hardware infrastructure that support low-and mixed-precision computation.The remainder of this report is organized as follows.Section 2 defines low and mixed precision.Section 3 defines energy metrics and tools to measure energy, and discusses the effects of programming languages in the energy consumption of code.Section 4 presents strategies, experiments, results, and analyses of mixed-precision training of neural networks.Section 5 discusses and presents experiments of mixed-precision strategies to solve symmetric positive definite linear systems, as well as its use in Gaussian Process Regression.Section 6 presents the final discussion and suggestions for future work. Mixed and Low precisionDue to finite memory, all computer software run at finite precision.The most well-known numeric types are the integers and the IEEE 754 single-precision (FP32) and double-precision floating-point (FP64) types.Integers are represented with 32 or 64 bits in modern 64-bit architectures, while FP32 and FP64 always have 32 and 64 bits, respectively.Based on this usual scenario, we consider low-precision types to be the ones whose size varies between 1 bit and 31 bits.Some examples of low-precision floating-point types are: TF32 (19 bits) 1 , FP16, BF16 2 , E4M3 (8 bits) and E5M2 (8 bits) [10].Two of the most used floating-point low-precision types are the BF16 and the FP16.Those are 16-bit data types where the mantissa has 7 bits in the former and 10 bits in the latter.They both use 1 bit to control the sign in front of the number.One attractive aspect of the BF16 type is that it can be easily converted to the FP32 data type since they both have the same number of bits in the exponent.Here is an example of a C++ class that emulates a BF16 and the cast operations to and from FP32.The cast operations use bit rotation, which is a common instruction implemented in hardware. struct bfloat162