A Semantic Driven CNN – LSTM Architecture for Personalised Image Caption Generation

L Abisha Anto Ignatious., S. R. Selva Jeevitha, M Madhurambigai., M. Hemalatha · 2019

Image Captioning is generating a human-readable textual description or a sentence about an image. The proposed semantic driven CNN-LSTM architecture comprises of the feature extraction process, semantic keywords extraction, facial recognition, and encoder-decoder LSTM networks. A pre-trained CNN is used to extract features from an image. A semantic keywords extraction module is used to identify the objects present in the image. The objects identified are labeled as the semantic tags present in the image. It increases the efficiency of captions in describing the objects and inclusion of these semantic labels in the captions. The LSTM based language model generates the captions by producing one word at a time. The facial recognition system identifies and recognizes the celebrity faces in the images, we have collected faces dataset which has facial images 232 celebrities. The instances of the person in the sentence were replaced with their names and personalized captions were generated. The Bilingual evaluation understudy (BLEU) and METEOR scores were generated to calculate the precision of generated captions.

Read the paper · More papers on PaperTik