Generating Descriptions for Visual Content Using Deep Learning Approach
Rajkumar B, Jaya Lakshmi A, P Deekshitha, V aacute rhelyi Tam aacute s, Krishna Dharavath · 2024
Image captioning using deep learning has achieved remarkable progress in generating human-like descriptions of visual content. This paper introduces a new method or approach that incorporates text-to-speech (TTS) functionality within a deep learning framework for live image captioning. The proposed system leverages Convolutional neural networks (CNNs) are used to extract features from images, while long short-term memory (LSTM) networks handle the caption generation. These generated captions are then input into a system TTS module, enabling the system to not only generate accurate textual descriptions but also convert them into spoken language. Providing real-time descriptive audio for visual content. This integration offers significant benefits by improving accessibility for visually impaired users and enhancing the user experience for sighted individuals. The paper details the system architecture, including pre-processing techniques for both live image and text data. We present the training process and discuss the evaluation metrics employed to assess the performance of the live image captioning and TTS components. Experimental outcomes on a benchmark dataset show the effectiveness of the proposed approach in generating accurate and natural language captions. Additionally, the integration of TTS functionality is evaluated through subjective user studies, highlighting its positive impact on user experience and accessibility.