PSDiffusion: Research on High-Resolution Facial Reconstruction Based on Diffusion Models
Xinhang Song, Qihan Bao, Nuo Cheng · 2024
With the rapid development of image processing technology, there is an increasing demand for customized, high-precision portrait graphics. The clarity and level of detail in these images directly affect tasks such as personalized design, virtual character creation, and facial beautification. Traditional methods for image customization are labor-intensive and time-consuming. Recently, diffusion models in the field of image processing have demonstrated significant progress in generating customized images. These computational approaches greatly reduce time costs while simultaneously providing multiple beautification and design options. To better generate art-enhanced images based on individual facial features, we propose the PSdiffusion (Portrait Styling Diffusion) model. To address the extended training times and high memory usage associated with traditional diffusion models, our model incorporates an Image Compression Module. Additionally, we optimize the Condition Information Module to enhance image controllability and mitigate the loss caused by compressed training. We further introduce an F-ESRGAN (Face-ESRGAN) high-definition module to improve image clarity, facilitating more detailed reference for designers. To address the lack of existing datasets for virtual character design, we also contribute a medium-sized, high-resolution dataset with detailed categorization and annotation. On the PSimage dataset, our model achieves an average MOS score of 4.23, with a PCT win ratio of 31.67 against three popular models, significantly outperforming most existing methods.