Enhancing Image Captioning Performance with VGG16 Feature Extraction and LSTM Sequence Processing
Dwika Lovitasari Yonia, Ikhsan Ariansyah, Aghnia Bella Dina, Nanik Suciati · 2024
This research explores the implementation and evaluation of an image captioning model using the Flickr8k dataset, which includes 8092 images with corresponding textual descriptions. Despite significant advancements in deep learning, many existing models struggle to consistently capture the intricate details and context within images, leading to less accurate and contextually relevant captions. This study addresses this gap by integrating VGG16 for feature extraction with LSTM for sequence processing, further enhanced by an attention mecha- nism that improves focus on relevant image parts. The model was trained on 6000 images and evaluated using the BLEU score, achieving a maximum score of 0.182. It also highlights areas requiring improvement, particularly in handling complex contexts and reducing computational demands. The findings sug- gest that although the current approach shows promise, further enhancements in computational efficiency and the adoption of advanced architectures, such as Transformer-based models, are necessary to achieve more robust and accurate image captioning performance. This research contributes to the ongoing development of image captioning systems by addressing key challenges in accuracy and efficiency.