A Structural-Aware Diffusion Generation Method for Traditional Patterns under Small-Sample Conditions

Quan Lei, Yidong Zhu, Jiaxin Cui, Wei Wang, Mengxi Niu · 2025

To address the challenges of style drift, detail loss, and mode collapse in traditional pattern generation under extremely low-shot conditions, this paper proposes a structure-aware optimization method for text-to-image diffusion models. Specifically, a multi-scale structure-guided adaptation module is integrated into the latent space of the diffusion model, jointly leveraging high-frequency detail reconstruction and class prior preservation losses. Additionally, a CLIP-based cross-modal semantic consistency constraint is introduced to ensure coherent alignment among textual descriptions, visual outputs, and geometric priors during the generation process. Using three representative datasets—blue-and-white porcelain, Dunhuang murals, and Miao batik patterns—each containing only 20 high-resolution (512 × 512) images, our method fine-tunes a limited number of newly added parameters while freezing the backbone weights. Experimental results show that the proposed approach reduces the Kernel Inception Distance (KID) from 6.5 to 2.3, improves LPIPS by 31%, SSIM by 4.5%, and CLIP similarity by 14.8%, significantly outperforming four state-of-the-art low-shot fine-tuning techniques. These findings demonstrate that the proposed model, via a synergistic mechanism combining structure guidance and multi-loss embedding space alignment, effectively balances content preservation and style adaptation. This method provides a promising and efficient pathway for the digital preservation and creative design of intangible cultural heritage.

Read the paper · More papers on PaperTik