Towards Efficient Learned Image Coding for Machines via Saliency-Driven Rate Allocation

Zixiang Zhang, Li Chen Yu, Qiang Zhang, Mei Junjun, Tao Guan · 2023

Recently, efficient image coding for machines method is widely required under machine vision task-oriented coding scenarios. To achieve higher task accuracy under a certain bitrate constraint, existing solutions take advantage of cutting-edge learned image compression methods and cascade the compression and specific task model straightly to carry out the joint optimization, which lacks adaptability and flexibility to handling machine vision saliency and non-saliency regions, thus obtaining inferior performance. In this paper, we propose a learned image compression model with a saliency-driven rate allocation strategy for machine vision. Specifically, an auxiliary network is introduced to provide reference latent first which contains full image reconstruction required information. And then, a machine vision saliency detector is applied to generate the adjustment factor to achieve rate allocation in the latent domain. In addition, a spatial dynamic convolution module is designed to facilitate the compressor adaptively processing the different spatial areas. Finally, the multi-distortion constraint is utilized to achieve an end-to-end optimization. Experimental results show that the proposed method achieves -77.48% coding gains over VVC using Mask R-CNN as the back-end task network on the Cityscapes dataset. Moreover, the reconstructed image maintains considerable visual fidelity, making it possible both for human and machine consumption.

Read the paper · More papers on PaperTik