Enhancing Landmark Detection Models Through Multimodal Fusion of Visual Data and Image Captioning

S.P. Raya, Imad Afyoni, Zaher Al Aghbari · 2025

Detecting specific landmarks in images remains a challenging task, especially when existing traditional vision-based detection models struggle to identify the landmarks accurately. In tourism, effective landmark detection is crucial for aiding navigation and enriching cultural exploration by assisting travelers in discovering points of interest, planning trip routes, and accessing detailed historical and cultural insights. This paper introduces a multimodal approach that enhances landmark detection by combining visual and textual features extracted from images. The model utilizes a deep learning-based feature extraction technique to generate the image embedding and employs an automatic image captioning technique to extract text embeddings. These embeddings are then concatenated into a unified multimodal representation vector for each image-caption pair. By fusing visual data with corresponding captions, the proposed method improves both the accuracy and robustness of landmark identification. The proposed method has practical benefits in tourism since it supports real-time use cases as tourist route planning and culture discovery. The model is trained on a custom-built Dubai landmarks dataset and tested in live image streams from Instagram, demonstrating its potential for real-world deployment. The proposed multimodal model outperforms the visual-based model by achieving an accuracy of 91.95%.

Read the paper · More papers on PaperTik