CPIG: Controlling the Portrait Image Generation by Distilling 3D GAN's Latent Directions
Ruiyan Wang, Jun Ling, Rong Xie, Li Song · IET Image Processing · 2025
ABSTRACT Synthesis of 3D‐aware facial images from latent spaces has garnered significant attention in multimedia content generation due to its ability to model images with rich semantics and diverse appearances. However, existing methods often rely on labeled data or suffer from incomplete attribute control and ambiguous latent space semantics. This paper proposes an efficient semantic distillation method that learns attribute directions of pre‐trained 3D GAN models, without the supervised semantic labels. We consider the latent space of GAN models as the mixture of two featured subspaces, namely the geometry‐aware space and appearance‐aware space. Following this hypothesis, we define two sets of learnable latent bases and use linear composition to represent controllable geometry and appearance feature space, respectively. To learn semantic‐wise latent bases for attribute‐controllable image generation, we design a framework and propose a three‐staged training strategy, which optimizes the appearance‐aware and the geometry‐aware latent bases. With the two sets of latent bases, we obtain the combined latent vectors using different weights for those bases and synthesize images with specified attributes. Compared to existing methods, our approach eliminates the need for labeled data and enables more controllable attribute disentanglement while ensuring identity consistency, which can be directly applied to real‐world scenarios such as virtual avatars and augmented reality applications. Experiments demonstrate the effectiveness and insight of our approach in aiding a better understanding of the latent space of 3D GANs.