From Diffusion to Decision: A Diffusion-ReRanking in Scene Text Detection
Yong Qiang JIA, Hong-Han Shuai, Hongxia Xie, Yung‐Hui Li, Hao‐Wen Cheng · 2025
Diffusion models have recently shown great potential in object detection and instance segmentation, yet their application to scene text detection, with its unique challenges such as instance variability and subjective human annotations, remains unexplored. In this paper, we propose DRR (Diffusion ReRanking), a method that adapts diffusion-based instance segmentation for scene text detection. Traditional instance segmentation often relies on classification scores for ranking, potentially overlooking the accuracy of bounding boxes and mask quality. DRR addresses this by incorporating two networks: a diffusion network, trained with a combination of projection loss and pairwise loss in the mask branch to produce more precise and tightly-bound segmentations, and a reranking network, which refines the results by evaluating bounding box accuracy and mask quality. Extensive experiments demonstrate the effectiveness of DRR, achieving a precision of 86.7%, recall of 81.7%, and an F-measure of 84.1% on CTW1500, highlighting DRR’s potential to advance scene text detection.