LR-DiffText: Low Resolution Scene Text Detection Based on Diffusion Probabilistic Model
Boyuan Chen, Zichen Dang · 2024
Recently, Diffusion Probabilistic Model (DPM) has demonstrated its remarkable performance in the field of image generation. The capability of DPM to efficiently simulate complex distributions and precisely capture image information has revealed their vast potential in image segmentation tasks. Scene text detection, as an important intersection of computer vision and text analysis, has garnered significant attention due to its extensive industrial applications. Particularly, segmentation-based approaches for variously shaped scene texts have become increasingly prevalent. However, existing scene text detection models generally rely on high resolution and high quality text images for training, which to some extent increase computational demands and limit the applicability of the model. We propose a diffusion probabilistic model for solving low resolution and irregular-shaped text detection, named LR-DiffText, which contains the following three features. Firstly, we propose a new dense conditional encoding structure that establishes dense vertical connections between encoding and the conditional encoding structures. Secondly, a concise and effective autoencoder architecture is introduced between the encoder and conditional encoder, ensuring efficient feature fusion and utilization within the conditional encoding component. Thirdly, significantly different from existing methods, LR-DiffText does not require high resolution scene text images for training, nor does it require massive datasets like SynthText for pre-training, thus greatly reducing the application difficulty and the computational load. The experimental results show that under the training condition of 256×256 low resolution data, the F-measures of LR-DiffText on Total-Text and SCUT-CTW1500 datasets can reach 79.1% and 84.8%, respectively. It is 4.4% and 8.9% higher than the highest F-measures of other popular scene text detection models such as DBNet at the same resolution. On lower resolution data of 128×128, LR-DiffText can still achieve F-measures of 75.3% and 78.4% on these two datasets, which proves the effectiveness of the proposed method.