Towards a Novel 8-Bit Floating-Point Format to Increase Robustness in Convolutional Neural Networks
Luis-J. Saiz Adalid, Juan-Carlos Ruiz-García, Joaquín Gracia-Morán, David de Andrés, J.-Carlos Baraza-Calvo, Daniel Gil-Tomás, Pedro-J. Gil-Vicente · 2025
Convolutional Neural Networks (CNNs) are widely adopted in Artificial Intelligence applications, particularly in computer vision and other deep learning tasks. Their performance relies on millions of parameters, including weights and biases, which are optimized during training, stored, and utilized during inference. Traditionally, these parameters are represented using the 32-bit IEEE-754 single-precision floating-point format. However, research has shown that excess precision in this format is not always required to maintain accuracy, motivating the adoption of reduced-precision 16-bit formats. A natural progression of this trend is representing real numbers using 8-bit formats. However, existing proposals often suffer from precision loss, negatively impacting CNN accuracy. In this extended abstract, we propose a novel 8-bit floating-point format designed to enhance reliability in CNNs, due to its reduced memory footprint, enough precision, and well-fitted range. We evaluate its advantages and limitations through comparative analysis. Initial findings suggest that our format improves computational efficiency while preserving accuracy comparable to 32-bit networks, increasing reliability. However, further experimental validation is necessary.