IDTS-IMAGE DESCRIPTION THROUGH SPEECH
International Research Journal of Modernization in Engineering Technology and Science · 2023
This work is a practical, workable way for converting image Text to Speech while keeping in mind the issues experienced by visually impaired people.The necessity for this study was spurred by the fact that engagement points for visually impaired persons are becoming more limited in a world that is becoming more digital, and that accessing digital media via an image describer can help these individuals.Visually challenged people can benefit from Deep Learning's ability to detect information contained in image objects.A review of the literature on image description technology introduces an alternative method that can automatically generate audio descriptions of images.This approach can substantially assist persons who are visually impaired.The processed visuals that the visually handicapped cannot see are then turned into appropriate descriptions and outputted as voice.Transformer Encoder-Decoder and Convolutional Neural Network (CNN) designs.The suggested approach creates a descriptive text caption for the image as an output after receiving an image and an audio input.The Transformer Encoder analyzes the audio input and creates a high-level representation of the audio features, while the CNN extracts the visual elements from the image.The text caption is then produced by the Transformer Decoder by paying attention to both the auditory and visual components.On the MSCOCO dataset, experimental results show that the proposed model outperforms current state-of-the-art models, achieving high accuracy and producing more evocative and coherent captions.The automated picture captioning feature of the proposed voice-based image caption generator may be used in assistive technologies.In essence, it scans the collected photographs for the presence of objects before converting them into userfriendly English.