ConvNeXt-ChARM: ConvNeXt-based Transform for Efficient Neural Image Compression

Ahmed Ghorbel, Wassim Hamidouche, Luce Morin · 2023

In recent years, neural image compression has garnered considerable attention from both research and industry. It has shown great promise in surpassing traditional methods in terms of rate-distortion performance through the development of end-to-end deep neural codecs. Despite these advancements, there is still room for improvement, particularly in reducing the coding rate while maintaining high reconstruction fidelity, especially in non-homogeneous textured image areas. Current models, including attention-based transform coding, also tend to have a higher number of parameters and longer decoding times. To address these challenges, we propose ConvNeXt-ChARM, an efficient ConvNeXt-based transform coding framework. It is coupled with a compute-efficient channel-wise auto-regressive prior that captures both global and local contexts from the hyper and quantized latent representations. Our architecture can be optimized end-to-end, fully leveraging context information to extract compact latent representations and achieve higher-quality image reconstructions. Experimental results conducted on four widely-used datasets demonstrate the effectiveness of ConvNeXt-ChARM. It consistently delivers significant BD-rate (PSNR) reductions, averaging 5.24% over the VVC reference encoder (VTM-18.0) and 1.22% over the state-of-the-art learned image compression method SwinT-ChARM. Additionally, we conduct model scaling studies to verify the computational efficiency of our approach. Furthermore, we perform objective and subjective analyses to highlight the performance gap between ConvNeXt, the next-generation ConvNet, and the Swin Transformer. Overall, our proposed ConvNeXt-ChARM framework showcases improved compression efficiency and reconstruction quality, establishing itself as a promising solution in the field of neural image compression.

Read the paper · More papers on PaperTik