Improving Image Captioning Using Deep Convolutional Neural Network
Fahimeh Basiri, Alireza Mohammadi, Ayman Amer, Mohammad Javad Ranjbar Naserabadi, Mohammad Mahdi Moghimi · 2023
Image captioning is a challenging task that requires a computer vision system to generate natural language descriptions for images. The aim is to build a model that can comprehend the visual content of an image and produce a coherent and accurate caption for it. This technology has various applications in domains such as assistive technology for visually impaired people, image search engines, image understanding in artificial intelligence systems, education, social media, and more. In recent years, recurrent neural networks (RNNs) with long-short-term memory (LSTM) units have been widely used for image captioning. However, RNNs have some limitations in capturing long-term dependencies and modeling complex language structures. In this paper, we propose a novel image captioning method that uses a convolutional neural network (CNN) to extract image features and a bidirectional encoder representation of transformers (BERT) to generate captions. BERT is an advanced language model that uses a transformer architecture to learn bidirectional context embeddings for words in a sentence. We train our hybrid model on a small-scale dataset of images and captions. We evaluate our model on several metrics and compare it with LSTM method. Our results show that our model achieves reliable performance and acceptable error rate (loss score). We also demonstrate that our model can generate more diverse and rich captions than RNN-based models.