Deep Fusion: A CNN-LSTM Image Caption Generator for Enhanced Visual Understanding
Chandradeep Bhatt, Sumit Rai, Rahul Singh Chauhan, Deepika Dua, Mukesh Kumar, Sanjay Sharma · 2023
Photo caption generators had gained huge recognition recently from the areas of NLP and computer vision. These generators produce captions that explain an image's content using deep learning algorithms. A sizable collection of image-caption pairs is used to train the algorithm. The photos are fed into a pre-trained CNN after the dataset has undergone preprocessing to extract its visual properties. Along with the picture features, the necessary captions are tokenized and supplied to the LSTM decoder. In order to accelerate the caption generation process, the LSTM decoder is trained to utilize beam search and maximum likelihood estimation methods. To assess the effectiveness of the suggested model, several metrics such as BLEU and METEOR are employed. The suggested CNN-LSTM image caption generator has a lot of promise for use in a variety of contexts, such as picture comprehension, information retrieval, and assistive technology for the blind. Combining CNN and LSTM models takes advantage of each architecture's capabilities, allowing the model to produce insightful and contextually appropriate captions for a variety of images. We also explore present difficulties and potential directions for additional investigation in this area. This research study concludes by presenting an innovative method for creating image captions that combines the strength of CNN and LSTM. The experimental outcomes demonstrate the efficacy and reliability of the suggested model, highlighting its potential to advance the field of image captioning and contribute to the creation of intelligent systems that can comprehend visual content and produce precise and insightful textual descriptions.