Smart Auto Image Captioning Using LSTM and Densenet Network
Arumugam M, Vijayakumar P, G M Siddharrth, J Rohini, Karthikeyan A, Chandrasekhar Rohith Bhat · 2023
The integration of semantic attention increase photo captioning ability. The premise underlying semantic attention-based techniques is that the model should prioritize words or qualities with high semantic value. Prior studies focused on the attribute detector and the captioning network as distinct entities. This hampers the best use of semantic data. Image captioning is the computational process of producing accurate written descriptions of an image's or video's visual content. Captioning images makes it easier to construct contextualized interpretations, allowing people to situate photographs within a larger context. Using photo captions can be useful in a variety of contexts, such as conducting research with an unlabeled image collection or attempting to identify. previously unknown patterns. Subtitles added to the photos using a Deep Learning Model. Recent advances in the fields of deep learning and natural language processing have made it much easier to generate captions for input photographs. This study will use neural networks to generate image captions on their own. The Long Short-Term Memory (LSTM) model, a type of recurrent neural network, used as an encoder, together with a pre-trained lexicon and visual attributes. This enables the extraction of image data and the generation of captions.