Fine-grained Medical Vision-Language Representation Learning for Radiology Report Generation

Siyuan Wang, Bo Peng, Yichao Liu, Peng Qi · 2023

Given the input radiology images, the objective of medical report generation is to produce accurate and comprehensive medical reports, which typically include multiple descriptive clinical sentences associated with different phenotypes.Most existing approaches have relied on a pretrained vision encoder to extract the visual representations of the images.In this study, we propose a phenotype-based contrastive learning framework, i.e., PhenotypeCLIP, to efficiently bridge the gap between visual and textual modalities for improved text generation.In contrast to existing contrastive learning methods which learn representations by contrasting images with entire reports, our approach learns more fine-grained representations, i.e., phenotype-based representations, by contrasting images with each sentence within the reports.The experiments on two widely-used datasets MIMIC-CXR and IU X-ray demonstrate that PhenotypeCLIP can achieve promising performances and substantially outperform the conventional contrastive learning methods.

Read the paper · More papers on PaperTik