TPG-DAN: A Text Prior Guided-Dual Attention Network for Scene Text Image Super-Resolution
Xinjie Feng, Wei Wang, Xiang Li · 2023
In recent years, scene text recognition (STR) tasks are usually affected by blurred and low-resolution text images, causing low recognition accuracy. A novel approach is to introduce scene text image super-resolution (STISR) as a preprocessing step to enhance text recognition performance. To further improve the super-resolution performance in text scenes, we propose a novel text-prior guided dual attention network (TPG-DAN) to reconstruct high-resolution text images, significantly improving text recognition accuracy. Precisely, we first extract the potential text representation in the image through the text prior (TP) module, enabling the network to focus more on the text area. Then, the sequence attention module (SAM) captures the contextual information between characters. This allows the network to correctly infer each character based on its front and back features. At the same time, the local attention module (LAM) is used to capture the characters' traits to recover compelling characters. We have conducted many experiments on the TextZoom benchmark, achieving text recognition accuracies of$79.62 \%, 64.28 \%$, and 45.87 % on the ASTER's easy, medium, and hard test sets, respectively. And evaluated the generalization of TPG-DAN on the ICDAR2015 and SVT datasets.