Advancements in Image Caption Bots: A Deep Learning Approach
Adarsh Dabral, Raj Gusain, Sushant Shekhar, Anurag Vidyarthi, Rıshı Prakash, R. Punitha Gowri · 2024
In recent years, artificial intelligence (AI) has changed the area of computer vision, and image captioning is one of the many applications that have benefited from these advancements. An image caption bot is an intelligent system that generates a textual description of an image. This research paper will explore the development of image caption bots, the underlying technologies, and the various techniques used to build these systems. Image caption bots are developed with the help of deep learning models (DLM) such as convolutional neural networks (CNNs) the most popular nowadays and recurrent neural networks (RNN s) which is also popular in researchers. Featureextraction has been done from the images with CNN sand RNN s are used to generate the captions. These deep-learning models have significantly improved the accuracy and performance of image captioning systems. During training, the CNN extracts the relevant features from the images, and the RNN generates the corresponding captions. In order to decrease the discrepancy between generated and ground truth captions, the model is trained. The accuracy and efficiency of image caption procedures could be greatly increased by using the image caption bots developed in this research. These systems have several applications and can greatly aid visually impaired individuals in understanding images, assist search engines to better index and search images, and improve the accessibility of social media platforms for users with disabilities. Further research in this field could lead to more accurate and comprehensive Image captioning systems.