Efficient Post-Training Quantization for FPGA-Based Neural Networks

Oumayma Bel Haj Salah, Seifeddine Messaoud, Mohamed Ali Hajjaji, Mohamed Atri, Noureddine Liouane · 2025

Recently, convolutional neural networks (CNNs) have shown remarkable performance in a variety of computer vision tasks. However, as CNNs become more complex, there is a higher demand for computational power, which typically necessitates the use of advanced hardware, limiting their scalability and broader applicability. This drives the need for optimization techniques to reduce these computational costs. We propose a uniform approach that enables the deployment of neural networks on FPGA platforms that cannot handle highprecision values. A common strategy to address this challenge is to perform low-precision computations through neural network quantization. The goal of this paper is to evaluate post-training integer quantization across different networks and criteria. We introduce an efficient framework designed to accelerate computations while simultaneously reducing both latency and memory overhead. This approach involves asymmetric quantization of the weight and activation matrices with the goal of minimizing these costs. Our experiments demonstrate a model size reduction of up to $75 \%$, with minimal accuracy degradation of less than $1 \%$ on benchmark datasets. Furthermore, we observe latency improvements of up to $3 \times$ compared to full-precision models. We believe that the proposed method will provide new insights into the interpretation of the quantization of neural networks.

Read the paper · More papers on PaperTik