A Unified Visual Saliency Model for Automatic Image Description Generation for General and Medical Images
Sreela Sreekumaran Pillai Remadevi Amma, Sumam Mary Idicula · Advances in Science Technology and Engineering Systems Journal · 2022
An enduring vision of Artificial Intelligence is to build robots that can recognize and learn the visual world and who can speak about it in natural language. Automatic image description generation is a demanding problem in Computer Vision and Natural Language Processing. The applications of image description generation systems are in biomedicine, military, commerce, digital libraries, education, and web searching.A description is needed to understand the semantics of the image.The main motive of the work is to generate description of the image using visually salient features.The encoder-decoder architecture with a visual attention mechanism for image description generation is implemented.The system uses a Densely connected convolutional neural network as an encoder and Bidirectional LSTM as a decoder.The visual attention mechanism is also incorporated in this work.The optimization of the caption is also done using a Cooperative game-theoretic search.Finally, an integrated framework for an automatic image description generation system is implemented.The performance of the system is measured using accuracy, loss, BLEU score and ROUGE.The grammatical correctness of the description is checked using a new evaluation measure called GCorrect.The system gives a the-state-of-art performance on the Flickr8k and ImageCLEF2019 challenge dataset.