SmartSight: Image Captioning-Driven Assistive Mobility System for Enhanced Scene Understanding

Jyoti Madake, Dhruvesh Kamble, Santosh Kandhare, Kapil Sangameshwar, Kritika Raina · 2025

SmartSight is an assistive mobility system designed for the visually impaired to achieve environment-specific scene recognition specifically in an Indian environment. At the center lies an image captioning module that uses EfficientNetB3 as feature extractor, followed by a Transformer-based decoder to create related captions with semantic accuracy. SmartSight integrates this captioning framework with embedded systems for the live processing of video and generates descriptive captions in real-time for each frame of video. The model is trained on a MSCOCO, and Flickr8K and an Indian environment dataset specifically designed for this purpose called IndiView, thus providing the model with strong performance across different environmental conditions. Testing on BLEU-1 (0.77), BLEU-2 (0.63), BLEU-3 (0.51), BLEU-4 (0.40), CIDEr (1.23) metrics identifies enhanced captioning accuracy compared to traditional RNN-based approaches. In terms of design for practical application, the system produces real-time captions converted to audio, thereby greatly enhancing the independence and mobility of individuals with visual impairments in dynamic and complex environments.

Read the paper · More papers on PaperTik