Automatic Indonesian Image Captioning using CNN and Transformer-Based Model Approach
Rifqi Mulyawan, Andi Sunyoto, Alva Hendi Muhammad · 2022 5th International Conference on Information and Communications Technology (ICOIACT) · 2022
Image captioning involves a phrase text generation or more for the visual content descriptions from images. Caption generation for an image is considered important to aid human activities in comprehending visual material, such as captions on medical images, human contact with robots, and helping visually impaired people explain visuals. Our study aims to create an Indonesian image description and assess how successful the caption's approach is. Translated Flickr8k datasets are provided for this investigation. This study employs two methods for captions generation from a photo: CNN with ResNet as the encoder, and Transformer, a self-attention-based mechanism as the decoder. Using distinct Indonesian datasets, named Flickr8k Bahasa, we used a Transformer-based approach to create an Indonesian captions generation model from photos. We demonstrated that our Indonesian Transformer-based strategy outperformed the old one, where the best results were obtained with BLEU-1 to 4, METEOR, ROUGE_L, CIDEr of 56.00, 41.17, 29.42, 20.57, 19.50, 44.16, 57.26, respectively. Besides comparing the model's performance using different CNN's pre-trained models, a larger CNN model did not guarantee any accuracy improvement through fifty epochs of training processes.