Generative Video Compression with a Transformer-Based Discriminator
Pengli Du, Ying Liu, Nam Ling, Yongxiong Ren, Lingzhi Liu · 2022
Deep learning has been successfully applied to image and video compression. Specifically, generative adversarial network (GAN) can compress images at low bit rates with sharp details and high perceptual quality. In this work, we propose a novel generative video compression (GVC) model with a transformer-based discriminator (TD), which learns non-local correlations within video frames to improve adversarial training. Besides, our GVC model incorporates a new loss to train the generator, which combines a base loss, a discriminator-dependent feature loss, and a perceptual loss. Experiments on HEVC test sequences demonstrate that the proposed GVC model provides superior performance at extremely low bit rates, compared to existing learned and traditional video coding schemes.