Effective Quantization Technique for Enhancing Model Performance on Resource-Constrained Devices

Nelum Andalib, Mennan Selimi · 2025

The deployment of deep learning models on edge devices is characterized by significant challenges due to limitations in computational power, memory and energy efficiency. In particular, quantization has evolved into a leading technique for reducing the size of the model and improving the speed of inference, while still mitigating the accuracy drop. This study analyzes the performance of quantization as an optimization method and explores the feasibility of training ML models directly on IoT devices. It focuses on the impact of quantization on two neural network models: LSTM (Long Short-Term Memory) and FNN (Feedforward Neural Network) using three schemes (tf, Float32, Int8). In a comparative analysis, throughput, memory consumption, latency and symmetric mean absolute percentage error (SMAPE) were used. The results show that Int8 quantization substantially reduces memory usage and improves latency and throughput, especially for the FNN model, which achieved 99,000 samples per second at a latency of just 0.01 ms. Despite its stable performance characteristics, the Int8 scheme provided benefits to the LSTM model. Deep learning models achieve better efficiency in real-time applications through Int8 quantization while preserving accuracy with minimal loss and gaining substantial performance benefits.

Read the paper · More papers on PaperTik