Lightweight Diffusion Network for Real-Time Style Transfer on Mobile Devices with Joint Image-Text Interaction
Fang Wang, Zhengdong Zhu, Jingyi Deng, Junting Liao · International Journal of Pattern Recognition and Artificial Intelligence · 2026
Aiming at the challenges of deploying diffusion models on mobile devices and the subjectivity of textual style descriptions, this paper proposes an end-to-end style transfer framework based on a lightweight diffusion network and joint image-text representation. A Contrastive Language-Image Pre-training (CLIP)-based cross-modal feature extraction scheme is designed to decouple style semantics and detail features from reference images, overcoming the ambiguity of pure text prompts. To enable real-time inference, a diffusion Generative Adversarial Network (GAN) hybrid architecture (UFOGen) is introduced to achieve single-step generation, replacing inefficient multi-step denoising. Furthermore, a lightweight network (FasterVAE, Faster Variational Autoencoder) is developed using separated convolution, transformer layers, key-value projection sharing, and Swish activation, significantly reducing parameters and computational cost. On a Xiaomi 14 Pro mobile device, the framework generates a [Formula: see text] stylized image in 1.45[Formula: see text]s. Experiments show that our method outperforms state-of-the-art approaches in SSIM, PSNR, and style loss. User studies also confirm its advantages in style accuracy, detail preservation, and color naturalness. This work provides a practical solution for real-time style transfer on resource-constrained platforms, advancing the deployment of diffusion models on mobile devices.