Dual-Mode Rounding Algorithms and Hardware for Posit-Based DNN Training: The Future of Mixed Precision Frameworks
Vishesh Mishra, Mahendra Rathor, Urbi Chatterjee · ACM Transactions on Embedded Computing Systems · 2025
The Posit number system provides a promising alternative to traditional floating-point (FP) formats for deep neural network (DNN) training by offering tapered precision and a wide dynamic range, addressing key limitations of conventional FP formats. While recent research has demonstrated the advantages of Posit-enabled training and inference for fixed-precision applications, the development of mixed-precision frameworks has been hindered by the absence of rounding algorithms for transitioning between Posit formats. This dependency has limited the practical adoption of Posits in DNN workflows. In this article, we present a Posit-based Mixed Precision Training and Inference (PMP) framework, leveraging Posit32, Posit16, and Posit8 for distinct computational stages. Posit32 ensures numerical stability in critical operations, Posit16 balances precision and efficiency for intermediate computations, and Posit8 significantly reduces memory usage during inference. Specifically, we introduce algorithms for converting Posit32 representations into Posit16 and Posit8 , and vice versa, under two rounding modes: deterministic and stochastic. Stochastic rounding is employed to mitigate precision loss in low-precision arithmetic. Furthermore, we propose a hardware-efficient Posit Multiply-Accumulate (pMAC) Unit that integrates deterministic and stochastic rounding modules, enabling efficient mixed-precision computations. We validate our framework on ResNet-18, ResNet-50, ResNet-152, MobileNet-v2, VGG-16, and EfficientNet-B7 (trained on ImageNet), YOLOv2 (trained on PASCAL VOC 2012), and BERT (trained on WikiText-2). Experimental results demonstrate up to 1.5× training speedup with Posit16 -based PMP framework and up to 6.5× training speedup with Posit8 -based PMP framework when compared with fixed-precision FP32 training, while maintaining comparable or superior accuracy. Moreover, hardware results show that the design overhead of integrating proposed deterministic and stochastic rounding modules with the pMAC unit is estimated to be around 4.6% only.