Leveraging Text Semantics for Enhanced Scene Text Image Super-Resolution
Li Chen, Jinsong Wu, Y. Y. Liu · Intelligent and Converged Networks · 2025
In recent years, due to the development of neural networks, super-resolution technology has made unprecedented progress. However, most existing super-resolution methods treat scene text images as normal images, ignoring the text information within them. This paper proposes to incorporate categorical priors specific to the text in scene text image super-resolution (STISR) model training process, called two-stage text prior super resolution (TTPSR) framework. The TTPSR framework first employs a parallel context attention network to restore the low-resolution image without incorporating text priors. Then, the obtained image is used for text recognition to obtain the text prior. The attention mechanism is then used to fuse the text prior and image features to guide the generation of the final high-resolution text image. Experiments have shown that the TTPSR model outperforms relevant state-of-the-art models in relevant metrics, peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and text accuracy, on the Textzoom dataset.