Learning True Rate-Distortion-Optimization for End-To-End Image Compression
Fabian Brand, Kristian Fischer, Alexander Kopte, André Kaup · 2022
Even though rate-distortion optimization is a crucial part of traditional image and video compression, not many approaches exist which transfer this concept to end-to-end-trained image compression. Most frameworks contain static compression and decompression models which are fixed after training, so efficient rate-distortion optimization is not possible. In a previous work, we proposed RDONet [1], which enables an RDO approach comparable to adaptive block partitioning in HEVC. In this paper, we enhance the training and boost the model performance by introducing low-complexity estimations of the RDO result into the training. It is well known that the setup during the training should be as close as possible to the setup during inference. Since including an RDO search into the training is computationally not feasible, we propose a fast variance-based criterion which we can use to approximate the RDO behavior during training. Additionally, we use the same criterion to propose a variance-adaptive RDO initialization which converges faster, needs fewer RDO passes. We can therefore decrease the inference runtime significantly. With our novel training method, we achieve average Bjøntegaard rate savings of 19.6% in MS-SSIM over the previous RDONet model [1], which equals rate savings of 27.3% over a comparable conventional deep image coder, similar to [2]. With our novel initialization method, we can reduce the number of RDO passes to one. Therefore, we need only half the time for RDO, while still saving 26.8% rate. When we do not perform an RDO search but instead only rely on the initial estimation, we still obtain remarkable rate-savings of 23.6%, needing no additional time for an RDO search. The full paper is available on arXiv [3].