Multimodal Framework Using CNN Architectures and GRU for Generating Image Description
Gali Tanishk Venkat Mahesh Babu, Selvani Deepthi Kavila, Rajesh Bandaru · 2022 2nd International Conference on Advance Computing and Innovative Technologies in Engineering (ICACITE) · 2022
Convolutional Neural Network (CNN) architectures and natural language processing have become vital in analyzing and generating text from images and videos in the present era. To attain image description in natural languages, and capture the semantic image information, the authors proposed a multi-modal framework that automatically describes an image's content. This multi-modal framework aims to create relevant natural language information about images and their regions using deep learning techniques (VGG16 and Xception over image regions and GRU for sentence development) and then convert the text into Google Text-to-Speech (GTTS), which will be used to produce speech as it will be helpful to visually impaired people. The authors used the MSCOCO-17 dataset and achieved the best accuracy of 80% i.e., 0.8 BLEU-1 score using the Xception model framework and 70% accuracy i.e., 0.7 BLEU-1 score using the VGG16 model framework.