Q-Net Compressor: Adaptive Quantization for Deep Learning on Resource-Constrained Devices
Manu Pratap Singh, Lopamudra Mohanty, Neha Gupta, Yash Bansal, Sejal Garg · 2024
Deep Neural Networks offers a high rate of accuracy, extensive, parameters, and intensive computational requirements that often characterize these networks. Due to high memory consumption and energy usage, this results in significant challenges when deploying DNN s on hardware-constrained devices. This paper introduces the “Q-Net Compressor,” a novel solution for installing substantial neural (Tao, Hou, et al. 2022) network training on constrained resources edge devices. Utilizing adaptive precision levels to maintain model functionality while reducing memory and computation demands. The Q-Net Compressor, implemented with TensorFlow and Keras, offers flexibility across diverse deep-learning architectures, striking a balance between significant model size reduction and the preservation of predictive power. This innovative approach contributes to the evolving deep learning model compression field with quantization, bridging the gap between computational demands and resource constraints. This research paper presents the implementation and evaluation of various quantization techniques on a Convolutional Neural Network (CNN) model trained using a COVID-19 dataset. The study aims to investigate the effects of quantization on model accuracy and size reduction. The CNN model is trained on a dataset containing COVID-19 images, and different quantization techniques are applied to the trained model. The techniques include post-training quantization, quantization-aware training, scalar quantization, vector quantization, centroid quantization, uniform quantization, non-uniform quantization and hybrid quantization methods. The results demonstrate the impact of each technique on the model's accuracy and size, providing insights into the trade-offs between model performance and computational efficiency.