Improving Image Captioning for Chinese Cultural Relics with Diffusion Language Models

Chenggang Mi, Yu Li · Journal on Computing and Cultural Heritage · 2026

Accurate and detailed image captioning is crucial for documenting and disseminating knowledge about Chinese cultural relics, yet this task is severely limited by its domain-specific nature and the acute scarcity of paired image-caption data. While paired visual-text data is limited, substantial volumes of domain texts about these relics often exist. We propose a novel framework for Chinese cultural relics image captioning that effectively leverages this abundant domain texts using diffusion language models (DLMs). Our approach involves pretraining a DLM on the large corpus of domain texts to instill domain-specific linguistic knowledge, followed by fine-tuning the pretrained DLM on the limited paired image-caption data, conditioned on visual features. Experiments demonstrate that this strategy significantly boosts captioning performance compared to methods that do not exploit the domain texts or use them less effectively. This work highlights the power of DLMs in leveraging readily available domain text to overcome data scarcity for complex vision-language generation tasks, offering a valuable tool for cultural heritage documentation and broader natural language processing applications.

Read the paper · More papers on PaperTik