Harnessing Approximate Computing for Machine Learning
Salar Shakibhamedan, Amin Aminifar, Luke Vassallo, Nima TaheriNejad · 2024
This paper explores the integration and application of Approximate Computing (AxC) approaches to Machine Learning (ML), especially Deep Learning (DL) models. We focus on four principal techniques-quantization, approximate multiplication, approximate in-memory computing, and input-dependent AxC. We demonstrate how each contributes to reducing the energy demands of current Artificial Intelligence (AI) systems, while maintaining acceptable levels of computational accuracy. These techniques may be deployed on software or hardware platforms. Quantization and input-dependent techniques can be implemented through software on general-purpose systems, enhancing flexibility and ease of deployment. Approximate multiplier and in-memory computing require specialized hardware integration, e.g., as custom System-on-Chip (SoC) or System-in-Package (SiP) solutions. We also discuss the crucial aspect of reliability, emphasizing robust design and error resilience to ensure the operational integrity of AI applications. By thoroughly examining these AxC techniques, the paper discusses an approach to designing energy-efficient and reliable AI accelerators, especially for SoC/SiP systems, providing essential support for use cases such as mobile and edge devices.