Scene Text Image Super-Resolution via Content Perceptual Loss and Criss-Cross Transformer Blocks

Rui Qin, Bin Wang · 2024

Text image super-resolution is a unique and vital task aimed at enhancing the readability of text images to humans. It frequently serves as a pre-processing step in scene text recognition. Nevertheless, due to the complex degradation in natural scenes, recovering high-resolution texts from low-resolution inputs is ambiguous and challenging. Predominantly, existing methods employ deep neural networks trained with pixel-wise losses, tailored for natural image reconstruction, yet neglecting the unique characteristics intrinsic to text. While a limited number of studies proposed content-based losses, these primarily concentrate on the accuracy of text recognizers, resulting in reconstructed images that may still be ambiguous to humans. Moreover, these approaches typically exhibit inadequate generalizability when dealing with cross-language cases. To this end, we present TATSR, a Text-Aware Text Super-Resolution framework, which effectively learns the unique text characteristics using Criss-Cross Transformer Blocks (CCTBs) and a novel Content Perceptual (CP) Loss. The CCTB, consisting of two orthogonal transformers, is designed to extract both vertical and horizontal content information from text images. The CP Loss supervises text reconstruction by integrating content semantics through multi-scale text recognition features, thereby embedding content awareness effectively into the framework. Extensive experiments on different language datasets demonstrate that TATSR outperforms state-of-the-art methods in terms of both recognition accuracy and human perception. Codes are released at https://github.com/Imalne/TATSR.git.

Read the paper · More papers on PaperTik