Food Text-to-Image Synthesis Using VQGAN and CLIP

Claudia Rachel Wijaya, Ida Bagus Kerthyayana Manuaba, Ardimas Andi Purwita · 2023

Image synthesis involves the generation of new images using techniques like generative models with a focus on achieving realistic and semantically aligned results. While existing models can effectively produce images representing singular objects like birds or flowers, they face challenges when it comes to generating realistic and semantically correlated images of complex food compositions. In this paper, we propose a novel application scenario leveraging the VQGAN + CLIP model to generate food images by utilizing the extensive RECIPEIM+ dataset. Our research objectives include investigating the effectiveness of the model in generating food and drink images based on targeted images and ingredient prompts. Furthermore, we evaluate the quality of the generated images using evaluation metrics such as the Frechet Inception Distance Score and Inception Score. Our results indicate that the performance of our proposed model is comparable to other similar models in terms of image quality and coherence, highlighting the potential of our approach in addressing the challenges of complex food composition synthesis. By identifying and addressing the limitations of our approach, we aim to advance the generation of diverse and visually appealing food images. Additionally, we suggest future directions for further research.

Read the paper · More papers on PaperTik