Error Analysis of Matrix Multiplication with Narrow Range Floating-Point Arithmetic

Théo Mary, Mantas Mikaitis · SIAM Journal on Scientific Computing · 2025

Abstract. High-performance computing hardware now supports many different floating-point formats, from 64 bits to only 4 bits. While the effects of reducing precision in numerical linear algebra computations have been extensively studied, some of these low precision formats also possess a very narrow range of representable values, meaning underflow and overflow are very likely. The goal of this article is to analyze the consequences of this narrow range on the accuracy of matrix multiplication. We describe a simple scaling that can prevent overflow while minimizing underflow. We carry out an error analysis to bound the underflow errors and show that they should remain dominated by the rounding errors in most practical scenarios. We also show that this conclusion remains true when multiword arithmetic is used. We perform extensive numerical experiments that confirm that the narrow range of low precision arithmetics should not significantly affect the accuracy of matrix multiplication, provided a suitable scaling is used. Reproducibility of computational results. This paper has been awarded the “SIAM Reproducibility Badge: Code and data available” as a recognition that the authors have followed reproducibility principles valued by SISC and the scientific computing community. Code and data that allow readers to reproduce the results in this paper are available at https://github.com/north-numerical-computing/narrow-range-FP-underflow-experiments and in the supplementary material ( narrow-range-FP-underflow-experiments-main.zip [224KB]). [Formula: see text]

Read the paper · More papers on PaperTik