Image Captioning with Convolutional Neural Networks and Autoencoder-Transformer Model

Selvani Deepthi Kavila, Moni Sushma Deep Kavila, Kanaka Raghu Sreerama, Sai Harsha Vardhan Pittada, Krishna Rupendra Singh, Badugu Samatha, Mahanty Rashmita · International Journal of experimental research and review · 2024

This study deals with emerging machine learning technologies, deep learning, and Transformers with autoencode-decode mechanisms for image captioning. This study is important to provide in-depth and detailed information about methodologies, algorithms and procedures involved in the task of captioning images. In this study, exploration and implementation of the most efficient technologies to produce relevant captions is done. This research aims to achieve a detailed understanding of image captioning using Transformers and convolutional neural networks, which can be achieved using various available algorithms. Methods and utilities used in this study are some of the predefined CNN models, COCO dataset, Transformers (enc-BERT,dec-GPT) and machine learning algorithms which are used for visualization and analysis in the area of model’s performance which would help to contribute to advancements in accuracy and effectiveness of image captioning models and technologies. The evaluation and comparison of metrics that are applied to the generated captions state the model's performance.

Read the paper · More papers on PaperTik