Improving Medical Image Captioning with a Context-Aware Knowledge Graph Transformer Framework
Aarti Sahitya, Shilpa Shinde · International Research Journal of Multidisciplinary Technovation · 2025
In this paper, we proposed a context-aware knowledge graph transformer framework for improving the caption of chest X-ray images. Normally the role of a radiologist is to interpret the chest X-ray or MRI image and write a detailed summary of finding patterns in a report. To generate an automatic detailed summary of the image the proposed framework is divided into three steps. The first step captures the visual feature of images using computer vision algorithms as Resnet 50 and Alexnet. The Second step uses the knowledge graph layer is employed for calculating the similarity between the tokens based on angel and token overlap to generate context-aware meaning of each token. The third step utilizes the transformer-based decoder to generate the detailed caption. The performance of the proposed model is compared against existing baselines including LSTM, CONV2D, and BI-LSTM architectures. The Proposed model outperforms baseline models by achieving higher evaluation scores in terms of evaluation metrics as 63% (BLEU-1), 61% (BLEU-4), 79% (RIBES), 85% (precision), 82% (recall), 82% (SPICE),and 79% (METEOR) demonstrating its effectiveness in medical text summarization.