Image Description Based Medical Image Retrieval

Prem Shanker Yadav, Himanshu Kumar Shukla, Chandrabhan Singh, Shudha Shukla, Upendra Kumar · Apple Academic Press eBooks · 2025

Recent developments in medical technology have led to a massive increase in the availability of medical diagnostic photographs. Designing a precise retrieval model to return the most relevant photos in response to the query is, therefore, a crucial area of urgent study. The complex content of medical pictures, however, cannot be sufficiently described by the visual descriptors. A novel image retrieval model based on picture captions is created to address this issue. The objective is to produce textual visuals for clearly identifying the relationship between things with captions. Here, the pretraining deep learning models are used to achieve picture captioning. Text length of caption, frequency-inverse document frequency (TF-IDF), and Bag of Words are hybrid to get the best match between picture captions and the query text. Bidirectional Long Short-Term Memory (Bi-LSTM), Convolutional Neural Network, Recurrent Neural Network (CNN+RNN), and Deep Hierarchical Encoder-Decoder Network (DHEDN) are used for image captioning on dataset created from Kvasir (a multi-class image dataset for gastrointestinal disease detection). Experiments show that DHEDN is outperforming compared to Bi-LSTM and CNN+RNN.

Read the paper · More papers on PaperTik