Medical Image Captioning using CvT and DistillGPT2

Konisha Kar, Shivam Verma Abhishek Nishad, Jayanti Rout, Ashutosh Soni, Surendra Kumar Nanda · 2024

The need for automation has increased substantially in the current era, where digitization plays an important role in every day's journey. As the healthcare industry increasingly harnesses the power of Artificial Intelligence (AI) and Computer Vision (CV), the ability to translate intricate medical images into descriptive, human-readable captions has emerged as a transformative innovation with multifaceted implications. This burgeoning field represents a convergence of advanced image analysis, natural language processing (NLP), and the imperative to enhance clinical decision-making, education, and communication among healthcare professionals. Medical image captioning can be particularly useful in providing chest X-ray (CXR) reports, which can help lessen the radiologists' workload and speed up patient treatment. However, the current image captioning models lacks the necessary diagnostic precision. In order to tackle this problem, we provide an innovative approach that achieves significant gains by using the Distilled Generative Transformer 2 (DistillGPT2) as the decoder and the Convolutional to Vision Transformer (CvT) as the encoder. Our suggested methodology produces reports that are more like to radiologists' reports on chest X-rays.

Read the paper · More papers on PaperTik