Design of Low-Cost and High-Accurate 8-bit Logarithmic Floating-Point Arithmetic Circuits
Botao Xiong, Xingqi Shao, Chang Liu, Shize Zhang, Yuchun Chang · IEEE Transactions on Very Large Scale Integration (VLSI) Systems · 2025
Recent studies suggest that the 8-bit floating-point (FP) format plays an important role in the deep learning, where the$E4M3$(4-bit exponent, 3-bit mantissa) is suited for the natural language processing model and the$E3M4$is better on computer vision task. In this brief, the logarithmic number system (LNS) is used to simplify the design of FP8 multipliers and dividers because the multiplication and division can be performed by the addition and subtraction in the logarithmic domain. Furthermore, this brief finds that the 3- and 4-bit logarithmic and anti-logarithmic (Antilog) converters can be effectively realized by {x,$x+1$} and {x,$x-1$}. As a result, compared to the standard$E4M3$and$E3M4$multipliers, the cell area can be reduced by 32% and 40%. Compared to the standard$E4M3$and$E3M4$divider, the cell area can be reduced by 61% and 67%. In addition, compared with the INT8-based design, the area of convolution core using proposed multiplier is reduced by 33%. The accuracy loss of the quantized ResNet-50, MobileNet, and ViT-B based on the proposed convolution core are −0.12%, +0.38%, and +0.8%, which are better than the INT8-based design. In the end, the proposed divider can be used in the image change detection. The false rate is slightly reduced from 2.97% to 2.95% compared to the standard$E3M4$divider.