Spatial-Temporal Fusion Network for Unsupervised Ultrasound Video Object Segmentation
Mei Wang, Dezhi Zheng, Qiao Pan, Dehua Chen, Jianwen Su · 2024
Automatic tracking and segmentation of lesions in ultrasound videos could assist in early diagnosis and treatment plan development. However, this task is quite challenging due to problems such as low visual saliency of the lesions and large variation between adjacent frames. In this paper, we develop a Spatial-Temporal Fusion Network (STFNet) for unsupervised ultrasound video object segmentation. First, an Edge Blur Enhancement Module is designed to extract and preserve the edge details of the target objects in ultrasound frames for spatial feature enhancement. Then, a Dynamic Alignment Module is developed to correct the inter-frame inconsistencies by aligning target objects from adjacent frames with those in the current frame for temporal feature enhancement. To incorporate both the spatial and temporal information, we further implement a mixed training strategy. These innovations collectively refine the model’s learning process and substantially boost segmentation accuracy. Extensive evaluations on real lymphoma ultrasound video data demonstrate the competitive segmentation results of STFNet. Specifically, compared with the best results among seven competing baselines, STFNet achieves the best scores in terms of region similarity ${\mathcal{J}}$, contour accuracy ${\mathcal{F}}$ as well as their average ${\mathcal{J}}\& {\mathcal{F}}$.