Image Captioning using Visual Attention and Detection Transformer Model
Yaswanth Eluri, N. Vinutha, M Jeevika., Sai Bhavya Sree N, Gangishetti Abhiram · 2024
Image caption generation has witnessed significant advancements with the integration of Deep Learning (DL) models. By leveraging DL techniques such as InceptionResNetV2 for feature extraction and transformer-based architectures for natural language processing, achieves remarkable results in generating descriptive captions for images. Unlike traditional Recurrent Neural Network approaches, which suffer from issues like vanishing gradients and lack of parallelization, this method offers improved efficiency and scalability. The synergy between DL models and Natural Language Models enables the system to capture intricate sematic relationships between visual content and textual descriptions, resulting in more accurate and contextually relevant captions. By integrating InceptionResNetV2 and Detection Transformer, this approach leverages the strengths of both architectures, achieving state-of-the-art performance in object detection tasks. Through joint training, the model learns to detecting images and labelling/captioning on up to 50% occluded images with Precision of 0.982, Recall of 0.931, F1 Score of 0.942 and Sensitivity of 0.892.)