TPWGAN: Wavelet-aware text prior guided super-resolution for scene text images

Shengkai Liu, Jun Miao, Yuanhua Qiao, Hainan Wang · Image and Vision Computing · 2025

Scene text image super-resolution (STISR) is crucial for improving the readability and recognition accuracy of low-resolution text images. Many previous methods have incorporated text prior information, such as character sequences or recognition features, into super-resolution frameworks. However, existing methods struggle to recover fine-grained text structures, often introducing artifacts or blurry edges due to insufficient high-frequency (HF) modeling and suboptimal use of text priors. Although some recent approaches incorporate wavelet-domain losses into the generator, they typically retain RGB-domain losses during adversarial training, limiting their ability to distinguish authentic text details from artifacts. To address this, we propose TPWGAN, a GAN-based STISR framework that introduces wavelet-domain losses in both the generator and discriminator. The generator is trained with fidelity losses on the HF wavelet subbands to enhance sensitivity to stroke-level variations, while the discriminator processes HF wavelet subbands fused with binary text region masks via a spatial attention mechanism, enabling semantically guided frequency-aware discrimination. Experiments on the TextZoom dataset and several real-world benchmarks show that TPWGAN achieves consistent improvements in visual quality and text recognition, particularly for challenging text instances with distortions or low resolution.

Read the paper · More papers on PaperTik