Model Caption Generator Using Visual Geometry, Residual, and Inception Architecture
Agus Nursikuwagus, Rinaldi Munir, Masayu Leylia Khodra · 2022
Image extraction is to be an essential task upon image classification. One of the challenges of these topics is an improves the extraction model. After that, they combined it with a recurrent neural network to generate a word fitting to an image's area. Interpretation of geology images takes a long time. It needs many geologists, especially in the description of the content of image rocks. Based on these problems, this study proposed a model that can be conducted for the geologist tasks. It enabled to make a caption for an image of geology rocks. The study uses VGG16, ResNet, and InceptionV3 concatenate to LSTM and word2vec that successfully captioned images of the foreground object like cars, people, animals, and many others. Even though the model can extract the image, the outcome does not align with the research objective. The study confirmed that the outcome has a value BLEU score of 4-gram of 0.367, 0.344, and 0.273, respectively. The study outcomes still mistake identifying objects of background and do not correctly caption relate to rock contents. The study concluded that the proposed new model is to be open challenges to achieve a result precisely to geologist descriptions.