Permuting Accumulation Order for Low-Precision Machine Learning

Ebby Samson, Tony X. Liu, Wayne W. Luk, George Anthony Constantinides · 2025

Quantization of weights and activations in neural networks is widely used to reduce data movement and the computational footprint of multipliers in arithmetic units. However, this increases the relative area contribution of adders. Most recent work in neural network quantization uses large floating-point accumulators due the large rounding and clipping errors incurred by smaller accumulators even if weights and activations are otherwise quantized to narrow data types. In this work, we propose a novel method of finding and applying permutations to weight and activation order in neural networks to reduce the error induced by small floating-point adders for the multiply accumulate (MAC) functions in matrix multiplications. Our method optimizes the order of accumulation with very low computational overhead by using a shared ideal order for a group of vectors instead of an ideal order for each vector, and using a static order rather than dynamically generating one at runtime. Our technique does not require quantization-aware training (QAT) or modification of weights, making it applicable to large language models (LLMs).

Read the paper · More papers on PaperTik