CAPTION GENERATION OF IMAGES USING CNN AND LSTM

Ummar Yousuf, Ravinder Pal Singh, Monika Mehra · International Journal of Innovative Research in Engineering & Management · 2022

The contents of a picture are automatically created in Artificial Intelligence (AI), which combines computer vision and natural language processing (NLP) (Natural Language Processing). It is developed a regenerative neuronal model. Computer vision and machine translation are required. This model is used to produce natural-sounding phrases that describe the picture. Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) are used in this model (RNN). The CNN is used to extract features from images, while the RNN is used to generate sentences. The model has been trained in such a manner that when an input image is provided to it, it creates captions that almost accurately describe the image. On various datasets, the model's accuracy, smoothness, and command of language learned from picture descriptions are assessed. These tests reveal that the model typically provides correct descriptions of the input image.

Read the paper · More papers on PaperTik