Towards Reliable 3D Face Reconstruction from Voice: Evaluating the Consistency between Regression Metrics and SSIM

Atyanta Nika Rumaksari, Risanuri Hidayat, Rudy Hartanto · 2025

Voice-driven 3D face reconstruction has emerged as a critical technology for biometric and human-computer interaction applications, yet validating reconstruction quality remains challenging due to divergences between geometric accuracy and perceptual plausibility. While regression metrics (MAE, MSE, RMSE,$\mathbf{R}^{\mathbf{2}}$, Pearson) quantify parameter-space errors, they often poorly correlate with human visual perception measured through SSIM. This study systematically investigates the consistency between these validation paradigms using a multimodal framework. We develop a comparative analysis pipeline processing from VoxCeleb2 through ECAPA-TDNN for 192D voice embedding extraction; Three deep architectures (MLP, VAE, Transformer) mapping to FLAME-DECA's 100D shape parameters; and Dual validation via regression metrics and rendered mesh SSIM. Statistical consistency is evaluated through correlations and concordance analysis across facial identities. Experimental results reveal critical insights: The Transformer achieves superior MAE (0.0162) but comparable SSIM (0.5434) to VAE (MAE=0.0196, SSIM=0.5432). Strong inverse MAE-SSIM correlation ($\rho=-0.92$) confirms error reduction improves visual similarity. High$\mathbf{R}^{\mathbf{2}}$-SSIM consistency (CCC=0.95) suggests variance preservation crucially impacts visual plausibility. These findings establish that reliable validation requires integrating parametric accuracy (MAE$} \mathbf{0. 5 4}$), and cross-metric consistency ($\rho>0.9$), providing a framework for developing perceptually-grounded reconstruction systems.

Read the paper · More papers on PaperTik