Automated Image Captioning and Speech Synthesis
Amruta Mankawade, N Pavitha, Amruta VPatil · 2023
An innovative system that creates natural language captions for supplied photos and transforms them to voice for increased accessibility is Automated Image Captioning and voice Synthesis. The suggested system extracts important information from the image and generates a relevant and coherent caption using cutting-edge technologies of deep learning like Convolutional Neural Networks & Recurrent Neural Networks. Text-to-speech synthesis techniques are then used to translate the generated captions to speech, allowing visually challenged users to access the material. The suggested approach was tested using typical benchmark datasets and yielded encouraging results. Because the system can create natural language captions and convert them to voice, it May be used in diverse applications, consisting of media accessibility & assistive technology.