An Enhanced CNN-RNN based Deep Learning Approach for Image Captioning
Venkata Nagaraju Thatha, B. VeeraSekharReddy, K. Ryveshma, G. Karthik, G. Shiva, Ch. Vikram · 2025
In the present days, Image Captioning plays a crucial role in transforming the visual content to textual format. It combines computer vision (to read the image) and natural language processing (to produce a natural-sounding caption). The current state is to improve the deep learning methods for creating image captions by combining CNN and RNN networks with a speech recognition system, using images from the Flickr30K collection that have several written descriptions. The evaluation of a visual description generation model that combines convolutional neural networks (CNNs) with recurrent neural networks (RNNs) is to be done. The CNN model incorporates [VGG16], which is used to extract visual information from images, such as object forms and textures; these features are then processed by a [LSTM], which generates descriptive captions when incorporated into an RNN model. After being trained to understand and convert visual content into spoken language, the model can also translate these captions into spoken English with the use of an embedded speech recognition system