Octopus: A Latent Diffusion Model for Enhanced Text-Driven Manipulation in Image Synthesis

M Nithin Skantha, B Meghadharsan, C Sri Vignesh, J Thiruselvan, Arti Anuragi · 2024

By combining user interactivity and precise control over image attributes, the applications of text-based image synthesis can be vastly expanded. Thanks to recent advancements in developing state-of-the-art text-to-image models, we can now generate images with high fidelity and diversity. While a plethora of research focuses on improving the fidelity and diversity of images generated using text prompts, less focus is given to controlling the attributes or characteristics of the generated images. On the other hand, there have been advancements in image editing through text prompts, but most of these models lack the understanding of spatial knowledge and precise control over the objects present in the image, which makes the edited images lose much of their characteristics. To address these challenges, this study proposes a three-step methodology to introduce our text-driven image editing conditional diffusion model called “Octopus,” which has precise control over image attributes. To bring in control over image attributes, this study uses a similar methodology proposed in a recent study, which uses rich text to control image attributes while exploring different formats, including bold, footnotes, font size, and font color. Then this model is used along with a fine-tuned Large Language Model (Gemma), to create a dataset of image editing examples to train Octopus, which, given an input image and edit instructions, generates the precise edited image during inference. This study demonstrates the potential of this approach with a variety of compelling results.

Read the paper · More papers on PaperTik