Research and Application of Image Captioning Based on the CNN-LSTM Method
Jingwen Liu, Mingxu Gao · 2023
Image captioning refers to the generation of concise descriptions for input images by computers. Serving as a bridge between computer vision and natural language processing, image captioning finds extensive applications in fields like accessibility assistance, intelligent search, and autonomous driving. In this paper, this research propose a CNN-LSTM encoder-decoder architecture. Building upon the pre-trained ResNet-101, this research modify the network structure to serve as an encoder for image feature extraction. Subsequently, this research introduce an LSTM decoder with a Soft-Attention module to generate image descriptions. Lastly, this research optimize hyperparameters using the Optuna framework. Our approach demonstrates satisfactory results on the Flicker30k dataset, achieving a Bleu-4 score of 0.187, cross-entropy loss of 6.15, and a top-5 accuracy of 72%. The experimental outcomes compellingly establish the effectiveness of this algorithm.