Object Detection-Driven Image Captioning: Integrating YOLO with Natural Language Processing
Sunil Battula, K. Nageswara Rao, Siva Kumar Pathuri · 2025
One important field of study that combines language processing and computer vision to produce descriptive text from images is image captioning, which uses deep learning and natural language processing. In order to improve picture captioning, this study investigates the use of the YOLO method, a quick and effective object detection model. By identifying important visual components, YOLO's real-time object identification enhances the precision and pertinence of produced captions. We improve sentence creation by combining YOLO with language models, guaranteeing syntactic and semantic correctness. The suggested classifier produces the best captioning predictions for photos. We also go into datasets, assessment measures, and NLP approaches applied to picture captioning jobs. This method enhances visual comprehension, allowing for the creation of logical and contextually relevant captions.