Enhanced Vision:Hybrid Deep Learning Based Model for Image Captioning and Audio Synthesis

J. Renees Jenisha, C. Priyadharsini · 2024

Image processing for the visually impaired is crucial. It empowers individual with visual impairments to perform daily tasks independently, contributing to improved quality of life, education, and employment opportunities. This paper tackles the challenge of image captioning with limited data by combining deep learning and supervised techniques like CNN and LSTM. the combination of Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) network involves training a computer to comprehend visual information in an image and generate descriptive captions. The proposed model employs a multi-layered architecture, combining the strengths of CNN s in feature extraction from images with the sequential learning capabilities of LSTMs to generate coherent and contextually relevant captions. The model utilizes the Google Text to Speech (gTTS) library for text-to-speech conversion. The findings suggest that among the tested models, the CNN- LSTM model achieved the highest predicted BLEU score of 95% for an image caption. Evaluation of the proposed model demonstrates its effectiveness in generating captions that align closely with human descriptions.

Read the paper · More papers on PaperTik