Text-Driven Medical Image Segmentation With Text-Free Inference via Knowledge Distillation

Pengyu Zhao, Yonghong Hou, Zhijun Yan, Shuwei Huo · IEEE Transactions on Instrumentation and Measurement · 2025

Medical image segmentation plays a pivotal role in ensuring accurate diagnosis. Traditional methods are predominantly monomodal, relying solely on image data. These image-only methods require large amounts of high-quality annotations to achieve accuracy. However, acquiring such annotations is time-consuming and labor intensive. Recently, multimodal methods have been proposed to address this challenge by incorporating text prompts. Despite their potential, these image-and-text methods require text prompts during inference, which is impractical in clinical settings, where pre-existing text prompts for infected areas are not available. To overcome these limitations, we propose a novel text-driven segmentation method that combines the strengths of both traditional monomodal and multimodal methods. The method aims to reduce reliance on annotations by integrating textual information during the training process, while enabling segmentation prediction to be performed solely based on image data during inference. The core of this approach is a teacher-student network architecture. The teacher network, a multimodal model, integrates image and text information at each decoding stage through designed cross-attention modules. Meanwhile, the student network, a monomodal model, mimics the teacher network to absorb the learned text information during the corresponding decoding stages. Once trained, the student network can independently perform inference using only image data, making it highly suitable for clinical practice. Extensive experiments on four publicly available datasets demonstrate that the proposed method achieves superior performance compared with both existing monomodal and multimodal methods.

Read the paper · More papers on PaperTik