3D-Aware Latent-Space Reenactment: Combining Expression Transfer and Semantic Editing
Paul Hinzer, Florian Barthel, Anna Hilsmann, Peter Eisert · 2025
The latent space of Generative Adversarial Networks (GANs) forms a continuous manifold where each point corresponds to a realistic image. Traversing linear paths within this space yields smooth transitions in image appearance, preserving realism without abrupt changes. For example, interpolating between latent codes of a closed-mouth and an open-mouth subject produces a natural animation of mouth movement. Building on this property, we propose 3D-LaSR, a novel method for reenacting 3D human heads using latent space paths, derived from a single monocular video. Our approach leverages GAN inversion to extract latent vectors from video frames, and supports expression transfer between identities by applying latent animation paths from one subject to another. Consequently, our method outputs dynamic sequences of latent vectors. Unlike decoder-based techniques that produce fixed geometry, 3D-LaSR can directly manipulate the scenes using established GAN editing techniques to modify attributes such as hairstyle, age, eyewear, or expression. The method uses a key frame inversion approach, where intermediate frames are synthesized through linear interpolation. This significantly reduces computational cost compared to methods that require per-frame inversion. We validate our approach through a comprehensive set of experiments, demonstrating high-quality video synthesis, effective expression transfer, and flexible editability. Finally, as we develop our method around state-of-the-art 3D Gaussian splatting GANs, our method is able to render in explicit 3D environments such as video engines or VR settings.