Alternating Greedy Schedules: Enabling Low-Bitwidth Accumulation of Dot Products in Neural Network Computations

Vikas Natesh, H. T. Kung · 2025

We present Alternating Greedy Scheduling (AGS), an algorithm for avoiding overflow, specifically transient overflow, during low-bitwidth accumulation of dot products in neural network computations. In conventional quantized (e.g., 8-bit) dot products, partial results are accumulated into wide (e.g., 32-bit) accumulators to avoid overflows when accumulating intermediate partial sums. However, such wide accumulators increase memory bandwidth usage and reduce energy efficiency. We show that iterative N:M pruning in floating point followed by quantization to 8 (or fewer) bits, and accumulation of partial products in an optimal order (via AGS) allows for accurate, compressed models with a large number of partial products that do not require wide accumulators. We design, analyze, and implement the AGS algorithm to eliminate accumulation overflows at inference time for several neural networks. Our method offers a 2.7x reduction in accumulator bitwidth while achieving model accuracy on par with floating-point baselines for multiple image classification tasks.

Read the paper · More papers on PaperTik