The attention mechanism for intelligent text generation of traditional cultural symbols in artistic design
Xi Xu, D. Z. Guo · Applied Soft Computing · 2026
Amid the digital wave, how to leverage artificial intelligence technologies to accurately interpret and reconstruct traditional cultural symbols has become a key issue in the field of artistic design. Existing image caption generation methods can identify the visual features of objects. But they often fail to capture the inherent cultural spirit and symbolic meanings when dealing with cultural symbols with profound historical connotations and abstract metaphors. This results in generated texts that remain at the level of superficial visual descriptions, lacking cultural depth. In response to this semantic gap, this study proposes an intelligent text generation model integrated with a cross-modal attention mechanism. The model constructs a dual-stream encoding architecture based on the Visual Geometry Group 19 (VGG19) and Bidirectional Encoder Representations from Transformers (BERT). By introducing multi-head self-attention and cross-modal attention modules, it achieves dynamic alignment between visual regions and cultural semantics. Different from traditional methods, the model can actively retrieve key visual clues from images according to the generated textual context. It accurately describes the physical forms of symbols and effectively interprets the cultural implications behind them. Experimental results show that this model achieves scores of 0.42 and 1.31 on the Bilingual Evaluation Understudy-4 (BLEU-4) and Consensus-based Image Description Evaluation (CIDEr) metrics, respectively, significantly outperforming benchmark models. It provides a new paradigm with both accuracy and cultural expressiveness for the digital inheritance and innovative design of traditional cultural symbols.