Leveraging Deep Learning for Automated Image Captioning to Improve Search Engine Indexing

Krishna Kumar Mohbey, Malika Acharya, Shaik Abdul Ahad · Cureus Journal of Computer Science. · 2026

Modern search engines should be able to retrieve images effectively, but traditional techniques mostly rely on manually generated metadata and textual annotations, which are prone to inconsistency and incompleteness. These constraints lower the retrieval accuracy and scalability. This paper presents an image captioning system based on deep learning to produce descriptive, contextually rich captions that facilitate better human image indexing and retrieval. The semantically meaningful captions generated by the proposed system reflect the content and context of a given image. User interaction is provided through a Flask-based web interface, and MongoDB is an extensible, scalable database for storing images and their generated captions. Evaluation of the system was conducted on benchmark image captioning datasets and other image collections to assess its real-world performance. The comparative experiment was conducted against traditional metadata-based retrieval methods using the standard measurement metrics of precision, recall, and mean reciprocal rank. Findings indicate that the caption-based system is superior to traditional techniques in delivering more relevant search results and achieving a higher retrieval accuracy, especially when manual annotations are low or absent. The study shows that deep learning-based image captioning is a viable and scalable alternative to the conventional indexing methods. The system can be used to improve performance in applications such as digital libraries, multimedia archives, and content-intensive web platforms by incorporating semantic understanding directly into the retrieval pipeline, thereby reducing manual input and improving the relevance of retrieved material.

Read the paper · More papers on PaperTik