Image Caption Generator using EfficientNet

Sai Vikram Patnaik, Rohi Mukka, Roy Devpreyo, Ankita Wadhawan · 2022 10th International Conference on Reliability, Infocom Technologies and Optimization (Trends and Future Directions) (ICRITO) · 2022

A human can give a quick glance at a picture or a scenery and will be able to describe that scene very accurately and comprehensively. Image Captioning is a well-known but also a challenging task in research. The reason being that it is the intersection between Computer Vision and Natural Language Processing which should be able to develop and provide us with accurate description of an image. In this paper, we try to implement EfficientNet - the most powerful CNN model which uses less parameters and gives us a better prediction. We use the Microsoft: Common Objects in Context (MSCOCO) Dataset which contains around 330K images with 1.5M object instances and 80 object categories. We evaluate the model using BLEU-4 score, which indicates how similar the candidate text is to the reference texts, with values closer to one representing more similar texts. Our proposed model shows better understanding of the image and provides us with desirable results.

Read the paper · More papers on PaperTik