Progress of image caption: modelling, datasets, and evaluation
Zhenyu Wan, Fan Wu, Ke Xin, Liqiong Zhang · 2021 International Conference on Computer Information Science and Artificial Intelligence (CISAI) · 2021
Image captioning is a popular topic in recent years. Researchers tried different methods to improve the performance of the image captioning system. In this paper, three methods, including the vision-based method, semantic-based method, and attention-based method, are introduced. These three methods aim to enhance the ability to describe images. In addition, we gathered, compared, and analysed various datasets from different articles in which those three types of methods are used. Analysing the data of three types of methods using datasets such as MSCOCO and Flicker 30K, we get the result that the semantic-based method is the most efficient.