Deep Learning-Based Automatic Creation of Image and Video Captions: A Review
Hanumanth Raju R K, Vinay Kumar S B, KEERTHANA H MAGADI -, Jayashree R, N. Mekala, Tharunlal PMS · 2024
A key component of generative intelligence is the integration of language and imagery. A lot of study has been done on image captioning, or giving images meaningful words to explain them. Typically, a vision encoder and a language model are employed to generate words that describe the visual content. With the addition of object regions, attributes, multi-modal connections, LSTM attentive methodologies, early fusion approaches like Bidirectional Encoder Representations from Transformers (BERT) and many other neural network model examples, captioning models have seen substantial advancements over time. This study identifies the most important technological advancements in architectures used for both picture and video captioning, provides a reference to the body of literature, and highlights new developments in a field that combines natural language processing and computer vision to maximize their complementary effects.