Fusion of Computer Vision and Natural Language Processing: Automated Image Caption Generation
Chandra Prakash Akku, Srushith Cheemalwar, Vasanth Chamakura, Shiva Shankar Gogu, K. Swaraja, N. Arun Vignesh · 2024
In this work the advancement in the field of natural language processing and computer vision, introduces a robust image captioning model that harnesses the capabilities of attention mechanisms. The model encompasses several core components, including the loading and preprocessing of datasets and the integration of MobileNetV3small. The model architecture is designed with a multi-layer transformer decoder, indicating self-attention, and cross-attention mechanisms for a deeper understanding of the image and text data. Training this model involves fine-tuning MobileNetV3small and optimizing the transformer-based decoder for accurate and context-aware caption generation. The utilization of MobileNetV3small and the Transformer-based decoder marks a significant advancement in image captioning, and the model holds promise for applications in automated image description and content generation. The model was tested on diverse datasets, as exemplified by its performance on both the Flicker8k and Conceptual Captions datasets. The training process involves a comprehensive pipeline, encompassing dataset preparation, image feature extraction, and tokenization, ensuring a holistic approach to model learning. The comprehensive result analysis highlighted the enhanced captioning results of the proposed method with CIDEr values of 37.20 and 25.40 on the Flickr8k dataset and conceptual captions dataset.