Attention Unveiled: Revolutionizing Image Captioning through Visual Attention

Teja Kolla, Harsh Kumar Vashisth, Manpreet Kaur · 2023

Image captioning models are a type of "Natural Language Processing" (NLP) models that are designed to generate textual descriptions of images. These models are trained on large datasets of images and captions, and use a combination of deep learning models and natural language processing techniques to generate accurate and informative captions. These models have proven to be effective in improving the quality of the generated captions. In recent years, researchers have also been exploring the use of reinforcement learning, adversarial training, and other techniques to improve the performance of image captioning models. Specifically, techniques like Reinforcement Learning from Human Feedback (RLHF), where human-provided captions guide model training, Generative Adversarial Networks (GANs), which generate captions through a competition between a generator and discriminator network, and Self-Critical Sequence Training (SCST), which optimizes model performance based on its own generated captions, have gained attention. These approaches aim to enhance the quality and relevance of captions generated by image captioning models. Visual attention models use a transformer and attention models typically consist of two main components: an image encoder that extracts features from the image and a language decoder that generates the textual description. This paper focuses on techniques that can be used for Image Captioning using visual attention models.

Read the paper · More papers on PaperTik