Object Detection-Driven Image Captioning: Integrating YOLO with Natural Language Processing

Sunil Battula, K. Nageswara Rao, Siva Kumar Pathuri · 2025

One important field of study that combines language processing and computer vision to produce descriptive text from images is image captioning, which uses deep learning and natural language processing. In order to improve picture captioning, this study investigates the use of the YOLO method, a quick and effective object detection model. By identifying important visual components, YOLO's real-time object identification enhances the precision and pertinence of produced captions. We improve sentence creation by combining YOLO with language models, guaranteeing syntactic and semantic correctness. The suggested classifier produces the best captioning predictions for photos. We also go into datasets, assessment measures, and NLP approaches applied to picture captioning jobs. This method enhances visual comprehension, allowing for the creation of logical and contextually relevant captions.

Read the paper · More papers on PaperTik