Implementation of Image Generative models using ImageGPT
Payel Dutta, Kaustuv Kunal · 2023
Motivated by advancements in unsupervised representation learning for natural language, this paper sought to investigate whether comparable models could also generate useful representations for images. Using a sequence Transformer, we trained the model to autonomously predict pixels. GPT-2 scale model achieved impressive results, displaying robust image representations. To achieve the desired outcome, we implemented a two-stage approach, involving pre-training and fine-tuning. During the pre-training stage, we experimented with both autoregressive and GPT objectives. Based on findings, it appears that generative image modeling remains a promising avenue for acquiring high-quality, unsupervised image representations. This study will come in useful whenever there is a need for image completion. Given any random half image, the model will try to predict completed version of the image in multiple forms, which can be extended later to facial recognition, object identification as well as it might find its use in healthcare domain.