A New Hybrid Assistive Module Towards Text Recognition and Speech Conversion System from Real Time Acquired Images to Visually Impaired Peoples

B. Gobika, S. Megha, V. Sivasakthi, K. Tharageswari · 2025

Identification and classification of objects containing text is an essential step to assist the visual impaired peoples towards analyzing the acquired image through camera. In existing, familiar object recognition approaches such as VGG-16, DenseNet-121, ResNet-18 and YOLO were employed to identify and classify the object present in the acquired images with optimal speed and greater accuracy. Further those model produces less mean squared error to the captured imaged by camera. Despite of several advantages, those models are time consuming and it still produces the moderate accuracy on the dense distributions of the image captured with multiple objects and different dimension of the text present in those object. In order to tackle those challenges, an a new hybrid assistive module is designed and represented as Text Inferred Yolo model as YOLOv10. YOLOv10 has been employed to detect and select the multiple objects containing text in the image to assist the blind people with accurate text recognition using Tesseract OCR engine. Initially camera captured image or gathered image from coco dataset is partitioned into residual blocks in grid format containing equal dimensions. Partitioned image cells will be employed to the bounding box approach to extract the features of the grid. The model identifies the text features of the object and will be represented in form of text feature map using the layers of convolution neural network. Further Linear Convolution Neural Network has been incorporated to produce the text labels for the feature map containing various sizes through its fully connected layer. Linear CNN also smoothens the text edges of the class labels on processing it through normalization layer to effectively categorize characters in the extracted regions of the image. Sigmoid activation unit has used to generate the class label for the max pooled data. Finally, the segmented text is exposed to the blind people with sound using speaker and google text to speech API. Experimental analysis is carried out on proposed model composing the text detection and text recognition module explains the performance in terms of recognition accuracy.

Read the paper · More papers on PaperTik