Architecture Optimizations for Improving Neural Image Compression Compute Complexity

Matthew J. Muckley, Marton Havasi, Jakob J. Verbeek, Karen Ullrich · 2025

Data compression is a critical component of today's digital infrastructure, in particular for storage and transmission of visual data over the internet. While end-to-end neural codecs have been shown to outperform hand-designed image and video codecs, their mainstream adoption is hindered by their high computational cost. In this paper we reassess standard design choices for neural image compression autoencoders, identifying two inefficiencies: the first is the substantial compute spent on high-resolution feature maps, and the second is methods of applying activations that lead to less expressive representations. We mitigate both issues and propose a new architecture, PatchMixer, that begins with approximately patchwise-independent encoding, followed by mixing layers, thus enabling compute savings. PatchMixer achieves a BD-rate within 5% of VVC on Kodak with 97 kMACs per pixel of decoding complexity and 200 kMACs per pixel of cumulative complexity, to our knowledge the lowest cumulative complexity for a neural codec within 5% of VVC.

Read the paper · More papers on PaperTik