FRMM-LSTM multimodal fusion based caption recurrent network for vision analysis

Harapriya Kar, P. Viswanathan, Bajra Panjar Mishra, K. S. Kuppusamy · 2025

People with visual impairments worldwide often face challenges in accessing information about different environments. Captions based on scene content, along with image descriptions, have been found to help address these difficulties. Combining natural language processing with computer vision plays a vital role in improving scene understanding. However, without proper captions, image descriptions can become inconsistent. To address this issue, this paper introduces a multiregional fusion-based FRMM-LSTM neural system that generates descriptive scene captions. The fully connected recurrent multimodal long short-term memory (LSTM) gates learn from each modality present in the image. The model operates in two distinct phases, one for image input and the other for language, followed by a shared fusion phase to learn the relationship between the two multimodal features. A fully connected LSTM layer processes the language component, while the fusion stages in FRMM are used for training. The language features are then mapped onto a shared vector space, pre-trained using FRMM and LSTM, to produce coherent captions. Additionally, Caption Crawler technology uses reverse image search to find captions from the Web and global repositories, making them accessible to users with screen readers. The interface for the fused model is implemented using the flask framework, a Python-based web development tool. This work aims primarily to assist visually impaired individuals by offering a region-specific approach to image captioning.

Read the paper · More papers on PaperTik