Voice-based Image Captioning using Inception-V3 Transfer Learning Model

Vaibhav Thalanki, R. Nagha Akshayaa, R. Krithika, Rokeya Begum Jothi · 2023

This study presents a deep learning model to serve as an image caption generator that generates descriptions or captions of the images in proper natural language sentences, which will then be read aloud by the text to speech translator. With the growing demand for tools like this in various fields such as assisting the visually impaired, self-driving vehicles, and virtual assistants. Hence, the development of such systems has become increasingly important. The proposed system utilizes a combination of Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) with attention models, specifically by using the Inception V3 model and a variant of RNN called Gated Recurrent Units (GRU).

Read the paper · More papers on PaperTik