DeformableIR: Text Image Super Resolution Using Deformable Attention Transformer*

Jie Liu, You Zhang, Zaisen Ye · 2024

Text image super-resolution technology is widely used in the preprocessing stage of tasks such as scene text recognition to improve the readability of text images to humans. In order to facilitate important tasks and applications such as OCR recognition in the future, this paper proposes a method to improve the super-resolution of text images. The specific research work is as follows: 1. A research on Text Image Super Resolution Algorithm (Research on Text Image Super Resolution Algorithm Based on Deformable Attention Transformer) is proposed to address the issues of high memory and computational costs caused by the use of dense attention in text image super-resolution, the influence of irrelevant parts beyond the region of interest on features, and text deformation and distortion. This algorithm mainly uses Deformable Attention Transformer as an important part of the backbone network to generate super-resolution (SR) images [1]. 2. This article proposes introducing layout analysis in the field of intelligent document processing, which can recognize common layout elements in document images, including text, titles, and other elements, into Text Image Layout Label Module (TILL). This article uses the Image Processing Challenge dataset and TextZoom dataset, which can recognize the layout of text images in the synthesized dataset. In subsequent operations, the scale of image super-resolution can be adjusted as needed, which can improve the accuracy and effectiveness of image super-resolution. 3. Finally, the stroke position perception module (SPPM) and stroke content perception module (SCPM) were introduced [2]. The stroke position perception module can reduce the problem of character adhesion in the image reconstructed by the super-resolution network of the text image and enhance the feature extraction performance of the model by introducing character position information guidance, Enable the model to better focus on stroke position information. The stroke content perception module can enhance the discrimination of easily confused characters at low resolution, and the design of this module supervises the text image super-resolution network to generate more distinguishable images of easily confused characters. 4. This article addresses the issues with both Chinese and English text images in the current dataset. During training, models were trained separately for both Chinese and English text images. During testing, input images were classified, with Chinese text images using the Chinese model and English text images using the English model. The advantage of this approach is that this article uses a lightweight VGG binary classification model to determine whether it is a Chinese or English text. Due to the easy task of classification, the network only consists of a few layers of convolution and a very small number of parameters. This can reduce a lot of unnecessary calculations, saving time and improving efficiency, as well as saving space and other issues.After a large number of experimental results, it has been proven that the Transformer based text image super-resolution algorithm studied in this paper has good super-resolution effect, with PSNR of 35.82 and MSSSIM of 0.95. Provided more accurate text images for subsequent text image recognition tasks.

Read the paper · More papers on PaperTik