Mixed-Precision Post-Training Quantization for Learned Image Compression

Jie Yu, Songping Mai, Peng Zhang, Yucheng Jiang, Jian Cheng · IEEE Internet of Things Journal · 2025

With the rapid development of neural networks, learned image compression (LIC) has surpassed traditional methods in the field of image compression. However, current LIC frameworks rely on floating-point operations. On the one hand, the use of conventional floating-point operations results in high computational costs, limiting their deployment in the Internet of Things (IoT) devices. On the other hand, when employing prior models, floating-point errors in prior calculations can cause cross-platform decoding failures. Neural network quantization addresses these issues by using low-bitwidth fixed-point, which reduces computational complexity and memory usage while eliminating inconsistencies from floating-point operations. Although quantizing the entire network to a uniform bitwidth simplifies hardware deployment, it can significantly impair coding efficiency, particularly in LIC networks at high bitrates, where certain layers are more sensitive to quantization noise, and we are first to provide both empirical and theoretical insights into the underlying reasons for these layers’ increased quantization difficulty. To address these challenges, we propose a mixed-precision post-training quantization (MP-PTQ) scheme that requires only a small calibration dataset without retraining, enabling rapid deployment. This method assesses the quantization sensitivity of each layer through rate-distortion loss analysis, thereby allocating the optimal bitwidth and quantization scale for each layer. Our method significantly improves the compression performance of quantized LIC models, achieving a Bjøntegaard-Delta rate (BD-rate) loss of less than 1% compared to their floating-point counterparts. Furthermore, by employing per-tensor activation quantization, we enable hardware-friendly acceleration—achieving approximately 3× speedup over floating-point models.

Read the paper · More papers on PaperTik