The Visual Assistant - Image-to-Speech Generator

Amrit Raj, Sanchita Purohit Ghosh, Bharat Gupta · 2023

In this paper, we presented an image-to-voice conversion model designed to assist drivers. The model can be integrated into vehicle navigation systems to analyze images of roads and signs captured by onboard cameras. It can then convert into voice instructions, providing real-time guidance to the driver. This can be particularly useful in unfamiliar areas or when visibility is limited. Image caption Generator is a well-known research field within the field of artificial intelligence that focuses on image comprehension in addition to providing a linguistic description of the image. To be able to produce sentences that are well formed requires a syntactic as well as a semantic understanding of the language. Additionally, our model can contribute to the development of intelligent transportation systems by providing an accessible means for visually impaired individuals to navigate public spaces. By converting visuals into voice instructions, the model ensures that individuals with visual impairments can independently and confidently travel through urban environments, enhancing their mobility and inclusivity.

Read the paper · More papers on PaperTik