A lightweight diffusion model for image generation based on improved MobileViT

Yifei Liu, Wenyun Sun · 2024

Aiming at the problem that the structural complexity of the backbone architecture of denoising diffusion probabilistic model (DDPM) leads to the inefficiency of the training of the model, a lightweight image generation method (MobileDiT) based on MobileViT is proposed, which takes MobileViT block as the backbone architecture of DDPM, and improves the computational efficiency while keeping the image quality by combining lightweight convolution with Transformer. The conditional mechanism for introducing conditional information into the model is also improved by replacing the traditional layer normalization in the Transformer block with adaptive layer normalization and initialising each block as a constant function, allowing the model to process conditional information more efficiently. The experiment results show that the proposed model reduces the FID-50K to 2.15 and improves the IS value compared with models such as Style Generative Adversarial Network (StyleGAN) and Ablative Diffusion Model (ADM). This shows that the proposed model can not only improve the computational efficiency, but also enhance the quality of the generated images.

Read the paper · More papers on PaperTik