Text to Image Latent Diffusion Model with Dreambooth Fine Tuning for Automobile Image Generation

Muhammad Fariz Sutedy, Nunung Nurul Qomariyah · 2022

This paper will explain the generated text to image using Latent Diffusion Model while combining it with a pre-training model called DreamBooth fine tuning. There have been many studies relating text to image using generative models (GAN's), VAE, and ART. This paper will also try to generate an automobile image dataset using the newest model called Latent Diffusion model. Also, experiment to try pre-train the existing LDM model using a fine tuning method that has never been done before. The use of fine tuning is to pre-train the model and also to handle a smaller scale of dataset that this research had and to reduce the possibility of overfitting. The aim will be to generate text to image by combining 2 existence models above and to create synthetic images based on the text input that is wanted. Later on, the mini user interface will be made to let the user make an input based on the brand of the automobile and 1 or 2 conditional such as color, location, and automobile position. Furthermore, the measure of MSE and RMSE score as our parameter to determine how good the model as there is no parameter from researcher in data scientist for what is a good MSE score is, this model that it could generate low MSE score that is score below 0.3 and RMSE below 0.6 where good RMSE score determined between 0.3-06.

Read the paper · More papers on PaperTik