Image Caption Generation using CNN and Audio Conversion
K. Sindhu Priya, R.Vijay Babu, M.Muralidhar Reddy, Trishala Reddy, M. Maanesh · 2024
This research focuses on developing a system that can generate descriptive captions for images and convert them into audio output. The system employs a Convolutional Neural Network (CNN) to extract visual features from images, which are then fed into a Gated Recurrent Unit (GRU) to generate textual descriptions. The generated captions are subsequently converted into audio using text-to-speech techniques. By training the model on a large dataset of image-caption pairs, the system learns to associate visual information with textual descriptions, enabling accurate and coherent caption generation. This research contributes to the advancement of image understanding and generation, with potential applications in various fields such as image search, accessibility, and content creation.