Text Retrieval from Images in Voice Form

Deepak Saini, Rajesh Prasad, Harsh Mishra, Abhishek Gupta, Khushi Vishwakarma, Adarsh Dubey · 2024

We aim to enhance the comprehension of visual content for individuals who are blind or have visual impairments. Our main objective is to deliver precise and enlightening descriptions of photographs. This technology is advantageous not just for visually impaired individuals, but also for robots and companies. To enhance the accuracy of object identification in photographs, we are employing an upgraded iteration of YOLO V5. Afterwards, we utilize an Xception V3 model to produce words based on the identified things. Consequently, the outcome includes both a written depiction and an audio recording available in multiple languages. By employing multiple sets of photos, we conducted a series of tests to evaluate the precision of our method, which yielded an impressive accuracy rate of 99.5%. It surpasses other procedures that have been tried in the past. Concurrently, we explored further computer program-driven methods for describing images. Specifically, our focus of the study is on a specific form of a network called a recurrent neural network. Furthermore, it evaluates the efficiency with which different models analyze and articulate visual representations. We experimented with the utilization of a model called CLIP assists in the depiction of images. Our method remains effective while being faster and more user-friendly compared to earlier approaches.

Read the paper · More papers on PaperTik