Analysis of Different Feature Extractors for Image Captioning Using Deep Learning
Srushti Shinde, Dhaval Hatzade, Shubhankar Unhale, Gaurav Marwal · 2022 3rd International Conference for Emerging Technology (INCET) · 2022
With the surge in technological innovations and advancements, data is considered to be the most important commodity in today’s world. The collection and dissemination of massive volumes of data have swamped our world today. This has given rise to problems in domains like searching, indexing, and value extraction. Hence image captioning rises as a solution to all these problems. The literature survey revealed that there is a lot of fog around which architectures work best for feature extraction in image captioning. To tackle this gap we’ve proposed a centralized analysis on which algorithms give the best results in terms of image captioning in this paper. We’ve taken algorithms like Inception, ResNet, Recurrent Neural Network (RNN), Convolutional Neural Network (CNN), Visual Geometry Group 16 (VGG16), etc. into account for our system. In this analysis, we take into consideration all the important aspects that affect image captioning to maximize the results at the end. Dataset taken into consideration is Flickr 8k. Hence the fundamentals of the approaches are discussed in this paper to examine their performances, strengths, and constraints.