Image caption generator using deep learning architecture with LSTM and their comparative analysis
Jasiul Rahaman H. Shaikh, Aditya K. Shinde, Faiz Rangari, Mohammed S. Mohsini, Amey S. Talekar · IET conference proceedings. · 2026
Access to precise visual information is vital for digital accessibility, particularly for blind and visuall y impaired individuals. Existing image captioning models still generate correct captions that make sense only when trained on tidy curated datasets, which is a constraint for real applications. The variability in this is problematic for assistive technolog ies and thus, consistent and reliable generation of captions is of critical importance. To overcome these issues, we used and compared three deep learning models VGG16, ResNet50 and EfficientNetB2 combined with an LSTM language model for automatic image captioning. From these, the EfficientNetB2 + LSTM outperformed the rest as it had the best optimized features and was the most efficient architecture that provided the best answer to caption reliability problem. They report BLEU-4 and ROUGE scores of 0.26 and 0.59, respectively, as well as a CIDEr score of 0.89 on the Flickr30k dataset. This demonstrates the objective performance of the model. On a subjective level, the captions produced were logical and contextually appropriate, and showed great promise as a means to improve digital access. Our findings indicate that EfficientNetB2 + LSTM can be successfully deployed in assistive technologies as it is the most efficient and effective model we have tested.