Image Captioning Engine: Using Deep Learning
Puneeta Singh, Vasu Bhatnagar, Sarthak Razdan, Subham Dash, Satvik Yadav, Charu Awasthi · 2025
Image captioning is a model development process has the ability to generate descriptive captions in an automatic format for images by learning visual features and mapping them to natural language sequences. In this task there are various techniques involved from the process of “computer vision” and “natural language processing”. It utilizes deep “neural networks” to connect the gap between visual presentation and textual description. “CNN (Convolutional neural network)” and “RNN (Recurrent Neural Networks)” are employed to extract spatial features from images and generate fluent, contextually appropriate captions respectively. More recently, Transformerbased architectures, such as those used in models like GPT and BERT, have enhanced the ability to produce more contextually relevant and fluent captions. Image captioning comprises with wide spectrum of applications, from assisting the “visually impaired” and improving content retrieval systems to enhancing human-computer interaction and automating social media analysis.