LDM-styler: a latent diffusion network for semantic-aware oil painting style transfer and cross-modal imagery reconstruction
Linchong Ji, Zhiyong Liu, Yuqi Wang · Journal of King Saud University - Computer and Information Sciences · 2026
Latent diffusion models have significantly advanced the field of artistic style transfer. Nevertheless, extant methods continue to be subject to content distortion when strong oil-painting stylisation is applied. Furthermore, the semantic-preserving style transfer and cross-modal text-to-image reconstruction are conventionally approached as discrete tasks, resulting in a lacuna in integrated controllable generation frameworks. The present paper puts forward LDM-Styler, a unified latent diffusion framework for semantic-aware oil painting style transfer and cross-modal imagery reconstruction. The proposed framework incorporates three primary components. Firstly, a Semantic-Style Coupling Module (SSCM) is employed, which utilises DINOv2 semantic features to guide region-aware style injection and mitigate content-style misalignment. Secondly, a Cross-Modal Latent Reconstructor (CMLR) is incorporated, enabling text-conditioned latent reconstruction for flexible cross-modal generation. Thirdly, an Adaptive Brushstroke Texture Controller (ABTC) is integrated, providing interpretable control over oil-painting brushstroke attributes, including direction, thickness, and impasto-like texture. Experiments on MS-COCO and WikiArt demonstrate that LDM-Styler outperforms eight representative baselines in terms of fidelity, artistic quality, semantic preservation, and cross-modal alignment. Specifically, LDM-Styler achieves an FID of 8.9, corresponding to a 20.5% relative improvement over the strongest baseline IP-Adapter (FID = 11.2), and an ArtFID of 11.8, corresponding to a 17.5% relative improvement over IP-Adapter (ArtFID = 14.3). Semantic preservation reaches an mIoU of 0.82, which is 0.07 absolute points higher than IP-Adapter (0.75). Cross-modal reconstruction achieves an AUC of 0.91, and a blind user study shows an 87% preference rate for LDM-Styler over IP-Adapter. In conclusion, LDM-Styler provides a coherent foundation for the generation of controllable artistic content, combining semantic robustness, painterly texture modelling, and cross-modal guidance.