Improved medical image captioning for chest X-rays using a hybrid VGG-ELECTRA model
Jitendra Joshi, J Julie Christina, L. Remegius Praveen Sahayaraj, V J Sharmila, Ashwin Balasubramanian · Artificial Intelligence in Medicine · 2024
Analyzing medical images to gain insights into a person's well-being is of utmost importance in recent years. A hybrid version of ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately) combined with VGG-16 is incorporated in the proposed methodology to learn information from different regions of the image and to interpret them in Natural Language Processing (NLP). The proposed architecture generates a wide description for the medical images and is trained to improve the accuracy for generating accurate descriptive sentences for a given training image. The experiment was validated using Chest X-ray data from the National Institute of Health (NIH) and the findings were evaluated based on BLEU-4 and ROUGE-4 F1 scores. The proposed architecture achieves a 0.54 BLEU-4 score as well as a 0.95 ROUGE-4 F1 score that seems to outperform existing models used for chest X-ray image captioning tasks.