4DM-LDDM: A Mesh Latent Diffusion Deformable Model for 4D Facial Generation

Baodong Wang, Zhen Wang, Junli Zhao, Zhenkuan Pan · IEEE Access · 2025

4D facial generation aims to generate time-series 3D facial models, which has broad applications in digital entertainment, virtual reality, the metaverse, etc. However, this task remains particularly challenging due to the complex high-dimensional irregular mesh structures and the necessity to maintain both temporal smoothness and identity consistency. To address these challenges, we propose 4DM-LDDM, a mesh latent diffusion deformable model that innovatively integrates optimal transport theory and cascaded diffusion deformation models. Firstly, we employ a LSA-Conv-based mesh autoencoder with Wasserstein distance constraint to encode 3D facial meshes into a latent space. This framework ensures effective preservation of facial features while avoiding the high-dimensional complexity of directly generating 4D facial sequences based on the original 3D mesh. Subsequently, a diffusion deformation model in the latent space is proposed, which cascades two diffusion modules and one deformation module. The model learns spatial deformation information by estimating target latent codes and generating deformation fields via a spatial transformation layer, enabling continuous transitions from source to target meshes. This framework supports both unconditional and label-guided 4D facial generation. Experimental results demonstrate that our proposed 4DM-LDDM achieves smooth frame-to-frame transitions and high-quality facial animation sequences.

Read the paper · More papers on PaperTik