An Enhanced Hybrid Deep Learning Model for Efficient Automatic Image Captioning
Eliyah Immanuel Thavaraj A, Sujitha Juliet Devaraj, Anila Sharon J · 2023
Image captioning is a field that overlaps natural language processing and computer vision from deep-learning algorithms. In order to provide a coherent description for high-level picture semantics, it is also essential to be able to analyze the characteristics and the relation between the objects. This paper focuses on developing an approach to creating a semantic image caption that makes maximal use of context and image cognition. To handle image captioning, Linking the detected objects with attention-enhanced deep-learning models is something that interests us significantly. The YOLOv3 is utilized to extract the image's features, and an extended Recurrent Neural Network, RNN (Long Short Term Memory, LSTM) with attention enrichment is then employed in order to construct the caption. We have utilized the attention mechanism to generate captions for images after considering the recognized items in the image scene. The COCO datasets are used for all four proposed variant models and the subset used for testing. To assess how well the various models performed, the semantic similarity analysis between produced descriptions and the actual image description is carried out. Finally, our results are compared to other traditional models.