Generative Face Video Compression Using Depth Estimation and Compressed Sensing

Hao Chang, Cheolkon Jung · 2025

Generative face video compression (GFVC) aims to achieve ultralow bitrate face video communication via deep generative models. As the representative GFVC work, the Compact Feature Temporal Evolution (CFTE) model compresses the bitstream for inter frames to 4×4, and shows outstanding coding efficiency. However, the feature extraction suffers from difficulty in distinguishing foreground face and background, resulting in erroneous face generation. In this paper, we propose a GFVC model using depth estimation and compressed sensing. Depth maps provide prior information on the spatial face structure, while compressed sensing significantly reduces redundancy by leveraging low-dimensional sparse representation. Thus, we utilize depth maps to retain face geometry and reduce background interference in feature extraction, and combine them with compressed sensing to reduce feature redundancy and enhance face generation. The proposed GFVC encoder combines depth estimation and compressed sensing to perform compact and precise extraction of face features. Compacter bitstream requires less bandwidth during transmission, while more precise bitstream generates more accurate face images during decoding. The proposed GFVC decoder has a dual-branch structure based on optical flow estimation and multi-scale feature fusion to capture motion information and enhance the face reconstruction quality. Experimental results show that the proposed GFVC model achieves average BD-rate gains of 77.04% and 64.43% on official test sequences in JVET, i.e. Class A VoxCeleb and Class B CFVQA datasets, over VTM22.2 in terms of DISTS and LPIPS metrics, respectively. The code is available at https://github.com/Changhaocup/Depth_Estimation_and_Compressed_Sensing-GFVC.

Read the paper · More papers on PaperTik