Image Captioning using Domain Ontology based Semantic Attention Approaches
J. Navin Chandar, Ganesh Kavitha · 2024
Deep Learning has spawned rich interest in a wide range of research areas involving text, image, and video formats. Traditional neural networks have limitations to handling such data types that were overcome with innovative research in deep net architectures. Image Captioning is one such application that has progressed leaps and bounds making a huge impact in fields like medical diagnosis, assistive technologies, and its likes. The task here is to precisely and accurately describe the image that is being presented. Researchers have built various models to handle this adopting advanced deep learning techniques. However, the key challenge to be noted is that practically all applications of image captioning are applied to a particular domain of interest, generation of naïve or generic captions will not suit much to make an impact. This gap is bridged by proposing a novel architecture called Onto-IC that factors semantic attention by modeling the domain specific ontology that is in consideration. The Onto-IC architecture is a two-stage design wherein the first stage focuses on extraction of image features from datasets containing images and ground truth captions. In the second stage, the captions are generated from images by incorporating domain ontology induced attention mechanisms using LSTM network. The captions used for training include both original and domain specific semantic ontology injected captions. This addresses the concern of generating captions that are more aligned towards the domain under purview. To validate the hypothesis, Flickr8k dataset was utilized for training and the generated model is used to compute and evaluate the image caption metrics. An ablation study conducted on Onto-IC has shown BLEU-1 scores greater than $\mathbf{8 0 \%}$ and the presence of domain specific elements in captions significantly enhances the usage in real life applications. In the concluding section of this paper prospective ideas to further extend the work are discussed.