Image Captioning in Real Time
Ankit Patil, Karishma Saudagar, Atul Maharnawar, Tejas Rangatwan, I. Priyadarshini · Zenodo (CERN European Organization for Nuclear Research) · 2022
The current development in Deep Learning based Machine Translation and Computer Vision have led to incredible Image Captioning models using advanced techniques like Deep Learning. Even if these models are very accurate, they often rely on the use of exorbitant computation hardware making it problematic to apply these models in real-time scenarios, where their actual uses can be noticed. This model uses a hybrid CN-NRNN model, where the CNN part of the model system uses the Xception model for transfer learning, and RNNs are widely used in language modeling. The Flickr8k dataset is used for real-time training and testing. RNN’s LSTM model is used to avoid problems with extinction or gradient explosion during the training phase.