Attention-based Deep Learning model for Indoor Object Recognition Framework for Visually Impaired Individuals

Komal Mahadeo Masal, Shripad S. Bhatlawande, Sachin Dattatraya Shingade · 2024

An enhanced indoor object identification structure for visually impaired individuals is utilized in this study. In order to create a more feature-rich training image, the RGB and depth pictures are initially collective using an Information Translation Module (ITM). The combined images are used to build feature maps using the Attentional Dense201 model. Using the help of RPNs and SIG Convs, the network is able to better understand spatial relationships. Acquisition of ROI data is accomplished through the use of perception-dependent ROI pooling. The network's fully linked layer then uses the inputs to detect objects. The hyperparameters of the framework are adjusted using an IGOA algorithm. The proposed framework tests and trains on the SUN RGB-D dataset. Several existing approaches are contrasted with the proposed model to assess the efficacy of the model. The proposed model has a testing accuracy of 96.35% and a frame processing time of 0.083 seconds, outperforming all previous models.

Read the paper · More papers on PaperTik