Beyond Integer Quantization: Approximate Arithmetic for Machine Learning Accelerators

James S. Tandon · 2023

There is an insatiable need for compute in AI/ML as the largest models in use today require thousands of GPU years to train. As these models increase in size, mitigating the power consumption while increasing compute performance is essential. Quantization to smaller number formats such as 8-bit floating point for training and 8-bit/4-bit integer for inference are common methods for reducing complexity, but simply reducing the number of bits is not the only method for approximation. We introduce a method for approximate computer arithmetic for AI/ML based on a hybrid floating point-logarithmic (FPLNS) system. Our arithmetic approximation algorithms are demonstrated to have 4.36x power reduction and 4.21x latency reduction for 32-bit floating point multiplication in exchange for a worst-case relative error of 12%. We further demonstrate that this relative error leads to negligible loss in several popular AI/ML algorithms due to their resilience to error. We finally demonstrate that power/performance benefits from FPLNS approximate arithmetic compound with quantization.

Read the paper · More papers on PaperTik