Modelling Complex Associations for Image Captioning Using Vision Transformers

Chavda Ritul, Nalini Sampath · 2024

Captioning an image is the method of writing a plain English the description of an image. This is a difficult assignment since it calls for the model to comprehend the image's information and to produce correct and fluid text. It has been demonstrated that a novel class of neural network called vision transformers is useful for picture captioning. Vision transformers are able to learn long-range dependencies in images, which is essential for generating accurate and fluent captions. In this project, will explore the use of vision transformers for modelling complex associations in images. The work will investigate how vision transformers can be used to generate captions that are more informative and comprehensive. The work will also explore how vision transformers can be used to generate captions that are more creative and engaging.The work is based on Flickr8k Dataset, which is a large dataset of images and captions. The vision transformer model will be trained on this dataset. The model is then assessed using a held-out test set. Compared to conventional picture captioning models, the vision transformer model will be able to provide captions that are more detailed and insightful. It will also be able to generate more creative and engaging captions. This project has implications for the development of new image captioning models. Vision transformers are a promising new approach for image captioning, and they have the potential to generate more accurate, fluent, informative, creative, and engaging captions.

Read the paper · More papers on PaperTik