Case Study IV
Tariq Mohammad Arif, Md Adilur Rahim · 2024
This chapter presents an image captioning case study problem using deep learning. It uses an open-source image dataset with captions describing the salient features of images. CNN and RNN models are combined in this case study for training on the dataset images and captions. The chapter describes how to tokenize captions and convert sentences into a format that can be processed by the model. During the model training and caption generation, it utilizes the attention mechanism, a crucial part of image captioning that helps it focus on different important parts of an image. The chapter also presents other fundamental aspects of model training in a practical scenario, such as setting default configurations (batch size, learning rate, or the number of training epochs, etc.) and defining the dataset to structure and handle it effectively. For ease of understanding, a similar program structure and variable definition are maintained as the image classification, object detection, and image segmentation case studies. The chapter covers fine-tuning procedures to improve model performance by using the training dataset differently. The trained image captioning model is also tested using outside images for generating image captions.