Efficient Deep Neural Network Training With a Novel 5.3-Bit Block Floating Point Data Format
Mohammad Hassani Sadi, Chirag Sudarshan, Sani Richard Nassif, Norbert Wehn · IEEE transactions on circuits and systems for artificial intelligence. · 2025
Low-bit-width data formats offer a promising solution for enhancing the energy efficiency of Deep Neural Network (DNN) training accelerators. In this work, we introduce a novel 5.3-bit data format that groups fixed-point values sharing a common exponent and scaling factor within a block of data. We propose a two-level logarithmic mantissa scaling method, providing a wide dynamic range for the Block Floating Point based data format and mitigating data loss during the conversion of DNN parameters into the proposed format. Evaluation results demonstrate that the proposed data format enables DNN training with an average of 5.3 bits, with minimal accuracy loss and no required modifications to the training process. Additionally, implementation results of proposed data format shows 1.48$\boldsymbol{\times}$reduction in memory footprint compared to 8-bit floating point formats, while addressing the lack of sufficient dynamic range for DNN training in current low-bit-width formats such as 8-bit and below.