3DMM-GAN: Multi-Modal Alignment With Adversarial Learning for Compositional 3D Human Image Synthesis
Tianyi Chen, Hongxin Fu, Hao Wang, Yan Huang, Si Wu, Yong Xu, Dapeng Oliver Wu · IEEE Transactions on Emerging Topics in Computational Intelligence · 2025
Current 3D-aware Generative Adversarial Networks (GANs) struggle to produce high-quality human images due to their limited ability to effectively integrate multi-modal information, restricting realism and semantic accuracy. To address these challenges, we propose 3DMM-GAN, a novel 3D-aware GAN framework specifically designed for compositional human image synthesis. Our method uniquely integrates multi-modal data, including textual descriptions and color distributions, into the adversarial training process. We introduce a CLIP-based text-modality discriminator with part-wise supervision to enhance fine-grained semantic alignment between generated body parts and their corresponding textual descriptions. Additionally, a histogram-based discriminator is employed to enforce global consistency in color distributions and textures. Furthermore, a Cross Modality Smoother module is proposed to facilitate coherent alignment among multi-modal features. Extensive experiments on the DeepFashion-Multimodal dataset demonstrate that our proposed framework significantly improves the realism, semantic consistency, and geometric fidelity of synthesized human images, outperforming existing state-of-the-art 3D GAN and diffusion-based methods. Our approach sets a new benchmark for realistic and detailed 3D human image synthesis, paving the way for future research and practical applications in digital human modeling and virtual reality.