Depth-Aware Fashion Synthesis via Manifold-Aligned Multi-Modal Diffusion
Ruoan Ren, Yuying Jiang, Zhengyang Lu, Jie Lin Sun · IEEE Access · 2025
Fashion synthesis faces challenges generating geometrically consistent garment imagery respecting three-dimensional structure and cross-modal alignment. Existing approaches operate on planar representations, failing to capture fabric draping physics essential for realistic garment-body interactions. We present Riemannian Depth-aware Multimodal Garment Designer (RDMGD), a framework integrating metric depth and differential geometry for enhanced synthesis. Our approach treats input modalities (text, pose, depth, segmentation, imagery) as Riemannian manifolds with intrinsic properties. The core innovation, Riemannian Manifold Multi-Modal Alignment (RMMA), performs fusion through geodesic distance minimization and parallel transport, preserving manifold structure while establishing coherent cross-modal correspondences. The depth-aware U-Net incorporates hierarchical geometric attention modulating activations based on surface normals and curvature tensors from monocular depth estimation. This enables garment synthesis exhibiting physically plausible draping and proper occlusions. Curvature regularization prevents manifold distortion, ensuring stable convergence. Evaluation on VITON-HD and Dress Code datasets demonstrates substantial improvements: 50.5% FID reduction and 7.4% CLIP-Score increase over state-of-the-art methods. Human evaluation confirms 73.4% preference for RDMGD outputs and validates the high-quality photo-realism.