Emotion- and lip-aware monocular 3D face capture with transformers
Yuqi Zhou · 2025
Recent methods on 3D face reconstruction from monocular input images have made some achievements. However, existing methods remain constrained due to the use of limited facial image datasets, which often result in insufficient accuracy in capturing holistic facial details such as facial geometry, dynamic expressions, eye gaze orientations, and nuanced jaw and lip movements. Furthermore, the inherent inductive biases of convolutional neural networks (CNNs) impede model convergence and scalability. While current approaches prioritize detailed facial shape reconstruction and affective expression modeling, they often neglect realistic lip motion, thereby limiting their applicability to talking head area. To overcome these limitations, we propose a novel framework which leverages a hybrid dataset comprising diverse high-resolution, expressive, synthetic, and multi-view facial images and videos. Our method introduces a transformer-based reconstructor to jointly optimize shape, expression, pose, and scene parameters through a two-branch architecture. Additionally, we design specialized loss functions and regularization terms to enhance the estimation of eye poses, emotional expressions, and lip synchronization dynamics. Extensive experiments demonstrate that these innovations enable our framework to outperform current methods in both reconstruction fidelity and motion realism.