Generating Captions from Images using Unsupervised MO-CNN in Deep Learning
G Sangar, V. Rajasekar · 2024
The creation of algorithms that can produce textual descriptions of the contents of photos has attracted increasing attention in recent years. Using computer vision techniques to recognize the items and actions seen in an image, this activity, known as image captioning, entails creating natural language phrases that precisely describe these visual features. In this study, we offer a novel method for image captioning that combines recurrent neural networks (RNNs) with convolutional neural networks (CNNs) to provide more accurate and varied textual descriptions of images. Our model performs at an advanced stage on multiple benchmark datasets after being trained on a sizable dataset of pictures and the captions that go with them. We also thoroughly examine our model's output and show that it can produce interesting and engrossing descriptions for a variety of photos. Overall, our method is a significant advancement in the field of image captioning and has a broad range of potential uses in fields including image search, content creation, and assistive technology for the blind.