Image Captioning with Pretrained Language Generators
Saketh Vishnubhatla, Nishant Sinha · 2020
We present a novel framework for image captioning combining scene graph models with pre-trained language models. This is in contrast to previous works, which largely rely on an encoder-decoder like architecture. In our experiments we use a two-stage pipeline: a) generating scene graphs from an image and b) using pre-trained language generators to obtain captions. Using scene graphs leads to more grounded captions and helps in exploiting the visual context between different objects in an image. By fine-tuning pretrained language models, we are able to reuse the vast, compressed knowledge in these models for image captioning.