DEVELOPMENT OF IMAGE CAPTION GENERATION HYBRID MODEL

Aigul Mimenbayeva, Rakhila Turebayeva, Assem Konurkhanova · DOAJ (DOAJ: Directory of Open Access Journals) · 2025

This study presents a hybrid model for image captioning using a VGG16 convolutional neural network (CNN) for feature extraction and a long short-term memory (LSTM) network for sequential text generation. The proposed architecture addresses the challenges of producing semantically rich and syntactically accurate signatures, especially in languages with limited training data. The model effectively bridges the semantic gap between visual and textual modalities by utilizing pre-trained weights and a robust encoding-decoding system. Experimental results on a dataset of road signs in Kazakhstan show a significant improvement in inscription quality as measured by BLEU and METEOR metrics. The model achieved a maximum METEOR score of 0.9985, indicating high semantic accuracy, and BLEU-1 and BLEU-2 scores of 0.67 and 0.64, respectively, highlighting the model's ability to generate relevant and coherent captions. These findings underscore the model's potential applications in multimodal systems and assistive technologies. Using a pre-trained CNN model (VGG16), we can efficiently encode visual information by extracting high-level features from images. This approach is particularly useful for tasks that require consideration of the semantics of images, such as road sign recognition. The second LSTM model, as a sequence-oriented architecture, is well-suited for text generation, as it effectively considers the context and previous words in a sequence. These models can be integrated into systems requiring the analysis and description of visual information, such as autonomous vehicles or driver assistance systems. In conclusion, the proposed model demonstrates high potential for image caption generation tasks, especially in resource-constrained environments and for specialized datasets.

Read the paper · More papers on PaperTik