Comparison of various CNN encoders for image captioning

S Veena, K S Ashwin, Prateek Gupta · Journal of Physics Conference Series · 2022

Abstract Image captioning is the ability of machines to study the features of an image and then give a textual description of that image. It uses computer vision to extract the features from the image and natural language processing for generating the caption. Image captioning requires both Image analysis and natural language processing. It is a very important task for further research on visual intelligence in line with human perception. In this project we will compare the different CNN pretrained models that are available like VGG16, RESNET50, Xception and INCEPTION and find out which performs better for image captioning.

Read the paper · More papers on PaperTik