Model of Improved Floating Point 32-bits Quantizer
Nikola Vučić, Zoran Perić, Aleksandra Ž. Jovanović · 2022 57th International Scientific Conference on Information, Communication and Energy Systems and Technologies (ICEST) · 2022
This paper gives a proposal of an Improved Floating Point 32-bits (IFP 32) representation format as an upgrade of the Floating Point 32-bits (FP 32) format defined by IEEE Standard for Floating-Point Arithmetic (IEEE 754). Both standardized and proposed models are regarded as quantization schemes. They correspond to the piecewise uniform quantizer, so important quantizer's parameters are given. This article considers the quantization of data statistically modeled by the Laplacian distribution. Signal-to-Quantization Noise Ratio (SQNR) for both solutions is given in a wide range of variance and compared. These models can be used in signal processing and neural networks. Since the proposed model provides better performances with the same total number of bits, it can be used in situations when higher SQNR is needed.