Synthesizing 3D Faces and Bodies from Text: A Stable Diffusion-based Fusion of DECA and PIFuHD
Mitul Joby, Nihal Chengappa P. A, Nilesh Ravichandran, Pranav Rao Rebala, Nazmin Begum · 2024
This paper presents an innovative approach for generating realistic 3D face and body meshes from textual descriptions, combining the strengths of stable diffusion with advanced models like DECA (Detailed Expression Capture and Animation) and PIFuHD (Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization). Dual Stable Diffusion pipelines generate images corresponding to textual attributes for faces and bodies. These images are then converted into 3D models- the DECA model processes the facial image for cresting a detailed mesh and texture, hence generating a 3D face, while the PIFuHD model constructs the 3D body model from the body image. This transformative approach ensures realistic and diverse character synthesis, matching specific features such as age, gender, and certain other physical traits. The effectiveness of the method is evident in its visual quality with potential applications in virtual environments and entertainment, hence setting a new standard in the realm of character design and generation.