Medical Image Captioning using Deep Learning: A Vision-Language Approach

Rohan Paul C, Esther Daniel, S. Seetha, S. Durga · 2025

Medical Image Captioning aims to obtain an accurate and textual description from medical images, supporting healthcare professionals in diagnostic processes and reducing manual workload. This paper presents an automated solution using a VisionEncoderDecoder model that combines Vision Transformers for the extraction of visual features and BERT for natural language processing. The proposed model was trained and evaluated on the Indiana University Chest X-ray dataset, the system can produce concise and accurate captions for radiological images. Using beam search for decoding, the model achieves a score of 0.65, emphasizing its capability to generate captions closely aligned with ground-truth reports. Future work includes dataset expansion and applications to other medical imaging modalities such as MRIs and CT scans.

Read the paper · More papers on PaperTik