Sketch Captioning Using LSTM and BiLSTM

Yeeshant Dahikar, Diptee Vishwanath Chikmurge, Sharmila Kharat · 2023

The process of creating meaningful captions for hand-drawn drawings, known as sketch captioning, presents a distinct problem since it is hindered by the absence of intricate visual information. This research introduces an innovative methodology that integrates a Convolutional Neural Network (CNN) as an encoder and a Long Short-Term Memory (LSTM) or Bidirectional LSTM (BiLSTM) as a decoder in order to generate captions for sketches. The primary function of the CNN encoder is to extract significant features from the sketch pictures, including both global and local visual information. By capitalizing on the hierarchical characteristics of Convolutional Neural Networks (CNNs), the encoder is able to proficiently capture and depict the underlying structure of the drawing, along with its prominent elements. The decoder, which is based on LSTM/BiLSTM architecture, is responsible for receiving the encoded sketch attributes and producing captions that are both coherent and contextually appropriate. LSTM networks have gained recognition for their proficiency in modeling sequential data, rendering them well-suited for capturing the temporal relationships inherent in the process of sketch captioning. The potential of BiLSTM networks is enhanced by integrating forward and backward contexts, thereby facilitating a more comprehensive comprehension of the sketch context. In order to train our model, we use a large-scale dataset consisting of paired sketch-caption instances. These examples consist of human-made captions added to drawings. The model parameters are optimized via a mix of supervised learning and sequence generating approaches. In the training phase, the decoder is conditioned on the encoded sketch characteristics, enabling it to provide captions that are in accordance with the provided drawings. In order to assess the efficacy of our suggested methodology, we conducted a comparative analysis with established sketch captioning techniques, using the well recognized assessment measure BLEU. A user research study was done in order to evaluate the quality and relevancy of the produced captions. This work aims to compare the accuracies of LSTM and BiLSTM models by using the bilingual evaluation understudy (BLEU) score. The assessment is conducted using the flickr8k and mscoco datasets.

Read the paper · More papers on PaperTik