Enhanced Cross-Modal Attention Control for AIGC-Driven Artistic Image Generation

Qiong Yao, Wenxiang Shi, Jianqiao Zhao · International Journal of Pattern Recognition and Artificial Intelligence · 2026

Artificial Intelligence Generated Content (AIGC)-based artistic creation suffers from content distortion and excessive stylization due to weak multimodal fusion. This study proposes an Enhanced Cross-modal Attention Control (ECAC) model for controllable artistic image generation. Built on the Diffusion Transformer (DiT) backbone, the model integrates three-branch attention mechanisms to fuse text semantics, style embeddings, and content features. A learnable modality fusion weight matrix and temporal style intensity control strategy are introduced to balance style expression and content fidelity. Experiments on the ArtBench-10 dataset show that ECAC achieves an FID of 15.82, LPIPS of 0.247, and average style loss of 0.159, outperforming CycleGAN, Diffusion UNet, and Mamba-ST. Ablation studies validate the effectiveness of cross-modal attention and dynamic fusion. This study aims to provide a competitive pattern recognition scheme for AIGC art generation.

Read the paper · More papers on PaperTik