From Templates to Transformers: A survey of Multimodal Image Captioning Decoders
Himanshu Sharma, Devanand Padha · 2023
Automated text-to-image (image synthesis) and image-to-text (image captioning) generation are two of the most challenging and cutting-edge fields of study in Computer Vision (CV) in conjunction with Natural Language Processing (NLP). The image-to-text synthesis, also known as Image Captioning (IC), has numerous applications in visual assistance, machine vision, remote security, healthcare, and remote sensing, among others. Since their origin, the IC frameworks have been constructed utilizing a two-subcomponent pipeline consisting of the visual feature extraction and natural language modeling subcomponents. The recent surge of interest in the applications of sequential deep learning models in machine translation has resulted in the development of efficient language modeling architectures. Deep sequential modeling architectures, such as Recurrent Neural Networks (RNN), Long Short Term Memory (LSTM), and Gated Recurrent Units (GRU), tackle the intricacies of the multi-modular space far better than existing machine learning-based translation approaches. The tremendous growth of deep language modeling architectures in IC necessitates a comprehensive review of its literature. In this survey, we conduct an exhaustive and analytical analysis of the language modeling architectures used in IC frameworks for caption generation, along with their training corpus datasets. We also identify a list of open research issues and potential research areas for future work.