Integrating CNN and LSTM for Natural Language Descriptions of Images
Utkarsh Prakhar, Varun Shukla, Shreyansh Shukla, B.K. Mishra · 2026
Image captioning is a simple task that is at the intersection of both computer vision and natural language processing models because they require to read and understand the visual information and write logical text with regard to the visual information. This paper presents an image caption generator that is formulated on the application of the visual feature extraction model with the help of the Convolutional Neural Networks (CNNs) and sequential language modeling model with the help of Long Short-Term Memory (LSTM) networks. The CNN component synthesizes significant semantic data on the input pictures but the LSTM decoder synthesizes captions having contextual interpretations with the aid of similar semantics. The experiments with the recommended framework on benchmark datasets containing MS COCO show that the results have a significant improvement in fluency in captions, semantic accuracy, and syntactic coherence. The experimental results highlight the possibility of CNN-LSTM model to be generalised to different image classes thereby eradicating the disparity between imagery thinking and language expression.