EfficientNet-based Image Captioning System
Poonam Bansal, Kiran Malik, Sharvan Kumar, Charu Singh · 2023
An individual might take an instant glimpse of an image or environment and express it quite precisely and effectively. Image captioning i.e., automatic labelling of images is a well-known complex problem in the area of image processing. The rationale for this is because the convergence of Natural Language Processing (NLP) and Computer Vision (CV) should be capable of developing and providing the mankind with an appropriate depiction of a picture. In this study, we attempt to develop an EfficientNet-based model that requires fewer characteristics and produces excellent predictions. We utilise the very famous and well-known MSCOCO Dataset i.e., Microsoft Common Objects in Context, which includes near about 330K pictures, 1.5M object instances, and 80 different classifications of objects. We use the BLEU-4 scoring for the analysis of the system, that specifies how close the predicted caption is to the reference caption. Having BLEU score nearer to 1 denotes more comparable captions. The suggested approach comprehends the picture efficiently and produces more intended outcomes.