Implementation of Image Caption Generation Using VGG16 and Resnet50
S. Nehan, A. Vijaya Lakshmi, C. Shambhavi, G Chandrakanth · 2024
One of the biggest problems in computer vision is getting a machine to automatically show an image's content along with a phrase in natural language. Using deep learning models, this research demonstrates two separate methods to generate picture captions. The first method employs VGG16 for feature extraction followed by training a Long Short-Term Memory (LSTM) model. In the second approach, ResNet50 is utilized for feature extraction, and a combined Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) model is trained for caption generation. Both VGG16 and ResNet50 are popular convolutional neural networks for extracting characteristics from images, while LSTM and CNN-LSTM are recurrent neural networks suitable for sequential data processing. By comparing these approaches, we evaluate their effectiveness in generating descriptive captions for images. Experimental results indicate the strengths and weaknesses of each method, providing insights into the interplay between feature extraction and captioning models in image understanding tasks.