Ensemble Learning-based Image Captioning with NLP for Context-Aware Descriptions
Pothala Mutyalu Naidu, Shreyas Sai, M. Naveen, Ch. Anil Kumar · 2025
Image captioning is a challenging task that requires an appropriate combination of computer vision and natural language processing to describe an image accurately and with context relevance. This paper presents an advanced method for captioning images using an ensemble learning setup along with an NLP-based fusion layer. The proposed system analyzes images from various perspectives by using multiple pre-trained models to capture diverse features, and the generated captions are then combined into a single coherent description using the fusion layer. The system’s efficiency is demonstrated using popular evaluation metrics such as BLEU, ROUGE, and METEOR, showing high accuracy and computational efficiency. This system is designed to work efficiently for a variety of image types, ensuring scalability and real-time application compatibility. The proposed method addresses common challenges in image captioning, including ambiguity, context preservation, and ensuring relevance across different datasets. Qualitative evaluation confirms the system’s ability to produce human-like captions, making it suitable for a wide range of applications such as assistive technologies, autonomous systems, and content generation.