Advancing Medical Image Captioning: A Dynamic Convolution Encoder-Decoder Network Approach for Automatic Generation of Descriptive Captions for Chest X-Ray Images
Tarun Jaiswal, Manju Pandey, Priyanka Tripathi · IETE Journal of Research · 2025
Automation of medical report generation has the potential to improve patient care by enhancing the quality of treatment. The chest X-ray report (CXR) contains crucial information about a patient's chest and lungs, including any abnormalities or injuries. It is crucial for diagnosing various chest conditions like pneumonia, lung cancer, fractures, and heart conditions. This automated report could reduce physician workload and improve efficiency by rapidly analyzing large volumes of CXR images. Automation ensures data presentation and interpretation consistency, reducing human errors and maintaining the quality of research results. However, the accuracy of CXR findings remains a challenge. A new research direction aims to develop hybrid approaches combining computer vision and natural language techniques. In medical image captioning, conventional models utilize an encoder, commonly a Convolutional Neural Network (CNN), to convert input images into a vector representation of fixed dimensions. A Recurrent Neural Network (RNN) serves as a decoder, conducting language modeling and producing target descriptions utilizing the encoded vector. Recent CNNs employ identical operations across all pixels, though not all pixels hold equal significance. To address these concerns and facilitate disease detection, we propose a novel method incorporating a dynamic convolution operation on the encoder to improve image encoding quality. Additionally, we employ an LSTM-based decoder for language modeling and an attention network to enhance system robustness. Experiments using the IU-Chest X-ray datasets (CXRD) revealed our approach outperforms others, achieving a score of 0.587 on BLEU-1, 0.271 on Meteor, 0.387 on Rough, and 0.405 on Cider, indicating superior performance.