My Emotion on your face: The use of Facial Keypoint Detection to preserve Emotions in Latent Space Editing
Jingrui He, A. Stephen McGough · 2025
Generative Adversarial Network approaches such as StyleGAN and StyleGAN2 provide two key benefits: the ability to generate photo-realistic face images and possessing a semantically structured latent space from which these images are created. Many approaches have emerged for editing images derived from vectors in the latent space of a pre-trained StyleGAN and StyleGAN2 models by identifying semantically meaningful directions (e.g., gender, age or hair color) in the latent space. By moving the vector in a specific direction, the ideal result would only change the target feature while preserving all the other features. Providing an ideal data augmentation approach for gesture research as it could be used to generate numerous image variations whilst keeping the facial gestures and expressions intact. However, the entanglement issue, where changing one feature inevitably affects other features, impacts the ability to preserve facial gestures and expressions. To address the entanglement issue, we propose the use of an addition to the loss function of a Facial Keypoint Detection model to restrict changes to the facial gesture. Our approach builds on top of an existing model and simply adds the proposed Human Face Landmark Detection (HFLD) loss, provided by a pre-trained Facial Keypoint Detection model, to the original loss function. We quantitatively and qualitatively evaluate the existing and our extended model, showing the effectiveness of our proposed loss term in addressing the entanglement issue and maintaining the facial gesture and expression. Our approach achieved up to 49% reduction in the change of emotion in our experiments. Moreover, we show the benefit of our approach by comparing with state-of-the-art models. By increasing the ability to preserve the facial gesture and expression during facial transformation, we present a way to create human face images with fixed expression but different appearances, making it a reliable data augmentation approach for Facial Gesture and Expression research.