Latent space manipulation of stable diffusion model for image generation
Bin Qi · Theoretical and Natural Science · 2023
The creation of the diffusion model was a piece of exciting news in the deep learning community. Millions of people are excited to try the new DALLE 2 AI artist and the state-of-the-art stable diffusion model. One of the greatest advantages of these AI artists is that they can apply many distinct artistic styles to the same image. Nonetheless, the image created using a specific artistic style sometimes seems too strong or too weak. The ability of adjusting the strength of artistic styles of the generated images enables users of this algorithm greater degrees of freedom in terms of creating images of their wishes. Moreover, the ability of adjusting the strength of artistic styles of the generated images can also serve an educational purpose from which users can understand artistic styles in concreate terms. This paper proposed a novel method of adjusting the level of artistic styles of the generated image which utilized the continuity of the latent space of the Contrastive Language-Image Pretraining (CLIP) encoding and calculate the augmenting latent vector. This novel method provides a greater level of tailored manipulability of the generated image.