Visual Insight: Deep Multilayer Fusion with Inception-Based LSTM for Descriptive Image Captioning
Rashid Khan, Bingding Huang · 2024
The image caption is a technology that aids us in comprehending the contents while employing machines to create descriptive text for an image. The captions are generated using Natural Language Processing (NLP) and Computer Vision (CV). When the descriptions contain a single word like “boy,” “cycle,” etc., the image captioning work is completed by combining the detection method with image captioning when one predicted region covers the entire image, such as a boy riding a bicycle. to combine the tasks of localization and description It is presently a current hot trend in deep learning development to use it to analyze visual information and write descriptive text. This paper presented a multilayer dense focus image captioning model. We used transfer learning techniques to adjust pre-trained image classification models and integrate them with long short-term memory network (LSTM) architectures to evaluate the performance of each of the combined frameworks. The variable length input is encoded into a fixed-dimensional vector, which is taken as the maximum length of the caption available mapped with the image, and the recurrent neural network (RNN) uses this representation to “decode” it to the desired output sentence. We experimented with the Flickr8k, Flickr30k, VizWiz, and MSCOCO datasets. According to the analysis of experimental data on evaluation criteria, the model described in this research can effectively accomplish image captions according to the analysis of experimental data. Its performance is better than classic image captionlna algorithms.